Skip to main content
Version: Latest
User ManualNeed configuration, code or deployment details? In the Technical Manual: AI DocumentsAI Data SourcesAI Web CrawlersAI File Sources

Give the AI Your Knowledge

On its own, an AI model only knows what it learned in general. To answer questions about your prices, policies, products or help articles, give it your content. The assistant then looks up the passages that match each question, answers from them, and lists them as numbered sources under the answer.

MenuArtificial Intelligence > Profiles (Knowledge tab); Artificial Intelligence > Data Sources; Artificial Intelligence > File Sources; Artificial Intelligence > Web Crawlers
PermissionManage AI profiles; Manage AI data sources; Manage file sources; Manage web crawlers
FeatureAI Documents for Profiles, AI Documents for Chat Sessions, AI File Sources, AI Web Crawlers, and an indexing feature such as AI Data Sources - Elasticsearch
Can't find it?

What you see depends on your role. If the menu, a button or a setting on this page is missing, ask your administrator to give you the permission listed above or to turn on the feature. Only administrators can turn features on and change site settings.

Choose how to add your knowledge​

Your contentUseWho keeps it up to date
A handful of files, such as a price list or a policy PDFDocuments on a profileYou, by uploading a new version
Files a user wants to ask about in one chatAttachments in a chatThe user, for that chat only
Content already on this site, such as articles or productsA data source of type Search Index ProfileThe site, automatically when content changes
An external search index or database tableA data source of type Elasticsearch, Azure AI Search or PostgreSQLThe site, when you sync
A folder of files on a server, an FTP site or an SFTP siteA file sourceThe site, on a schedule
A public website, such as your help centerA web crawlerThe site, on a schedule
A public documentation site, searched only when neededA documentation search tool instanceNobody: it searches the live site

Before you start​

Knowledge is stored in a search index that understands meaning, not just words. Your technical team sets up the search service once. Then an administrator creates the indexes in Search > Indexing with Add index, choosing the type and an Embedding deployment (see Connections and deployments):

Index typeUsed for
AI Documents (Azure AI Search or Elasticsearch)Documents on profiles and attachments in chats. Then choose it under Settings > Artificial Intelligence > Documents > Index profile.
AI Knowledge Base Index (Azure AI Search or Elasticsearch)Data sources, file sources and web crawlers.

The embedding deployment cannot be changed after the index is created. Avoid switching to a different index once people use it, because documents stored in the old one are no longer found.

Upload documents to a profile​

Documents on a profile are background knowledge the assistant uses in every chat.

  1. Open Artificial Intelligence > Profiles, edit the profile, and open the Knowledge tab.
  2. Drag files onto Drag and drop files here, or click Browse Files. New files show a Pending badge.
  3. Click Save. The files are read, split into passages and indexed.

To remove a document, click its Remove document icon and save. Download document gets the file back.

Supported formats under the upload area lists the file types you can use. Usually that is text, Markdown, JSON, XML, HTML, YAML and log files, plus PDF, Word (.docx) and PowerPoint (.pptx) when those features are on. Spreadsheets (CSV and Excel) cannot be uploaded to a profile, because they are meant for calculations rather than for looking up passages. Old Office formats (.doc, .xls, .ppt) are not supported; save them in the new format first.

FieldWhat it does
Document Top NHow many matching passages the assistant reads for each question, from 1 to 20. Default 3. More passages give more context, but cost more and slow the answer.
Document retrieval modeChunk gives the assistant only the matching passages. Hierarchical finds the matching passages, then gives it the whole documents they come from: better for short documents, much more text for long ones. Default setting uses the site's choice.

The assistant uses the documents quietly; it does not tell users that it has attached files unless they ask about sources.

Let users attach files​

To let people attach documents or pictures to a chat, open the profile's Knowledge tab and tick:

FieldWhat it does
Allow session document uploadsUsers can attach documents to chats with this profile.
Allow session image uploadsUsers can attach pictures. Needs a default vision deployment.
Largest document that can be indexedThe longest document, in characters of text, a user can attach. Bigger documents are refused when attached. Use site default keeps the site's limit; 0 means no limit.
Describe figures in uploaded documentsWhether the AI describes charts and pictures inside attached documents, so it can answer about them. This uses a vision model and costs more. Use site default, Describe figures or Do not describe figures.

Attached files belong to one chat only. How users attach them is described in Chat with an AI assistant.

Data sources​

A data source connects the assistant to a knowledge base built from a search index or a database. Use it for large or changing content, such as all the articles on your site.

Create a data source​

  1. Open Artificial Intelligence > Data Sources and click Add Data Source.
  2. In Available Source Types, click Add on the type you need:
    • Search Index Profile: an index of this site's own content.
    • Elasticsearch or Azure AI Search: an external index.
    • PostgreSQL: a database table.
    • File: the target of a file source.
    • Web: the target of a web crawler.
  3. Enter a Name and choose the Destination index, the AI Knowledge Base index that stores the passages. It cannot be changed after you save.
  4. Fill in the source's own section, such as the Source index, or the connection details your technical team gives you for an external index or database. For an external source, Use the globally configured connection uses the connection your technical team set up.
  5. For every type except File and Web, fill in the Field Mapping: the Content field to search, an optional Title field for the source list, and a Key field that identifies each entry. Ask your technical team which fields to use.
  6. Click Save.

The data source fills its knowledge base and keeps it in step with the source. To rebuild it by hand, open the data source's Actions menu and choose Sync index.

Use a data source in a profile​

On the profile's Knowledge tab:

FieldWhat it does
Data sourceThe data source to answer from. No data source for none.
FilterSearches only some entries, for example only active products. Ask your technical team for the expression.
Restrict answers to retrieved data onlyThe assistant answers only from the data source. When nothing matches, it says so instead of answering from general knowledge. Turn it on for a public assistant that must stay on topic.
StrictnessFrom 1 to 5. How closely a passage must match the question. Higher values use fewer, more relevant passages. Empty uses the site default (3).
Retrieved documentsFrom 3 to 20. How many passages are retrieved for each question. Empty uses the site default.

The site defaults are under Settings > Artificial Intelligence > Data Sources: Default strictness and Default top documents.

File sources​

A file source reads a folder of files on a schedule and keeps a File data source up to date, so nobody has to upload anything. Use it for documents another system exports, such as nightly product sheets on an FTP site.

  1. Create a data source of type File (see Create a data source).
  2. Open Artificial Intelligence > File Sources and click Add File Source.
  3. In Available Connectors, click Add on File system (a folder on the site's server), FTP / FTPS or SFTP.
  4. Fill in the fields below and the connection details, then click Save.
FieldWhat it does
NameThe name shown in the list.
Target File data sourceThe File data source that receives the content.
EnabledWhether the scheduled reading runs.
Re-read interval (minutes)How often the folder is read again. The site checks once an hour, so shorter intervals still run at most hourly. Empty uses the site default.
FiguresAuto (describe charts and pictures that look worth it), All, or Off. Describing figures needs a vision model and costs more.
Figures described per documentThe most figures described in one document.
Vision deployment / Utility deploymentThe models used to read the files. Leave them on the site defaults.
Items per runThe most files read in one run.
LanguageThe language of the files, such as en-US, when they all use one.
Folder, Include sub-folders, Files per listingWhich folder to read. For File system, the folder is inside the area your technical team set aside for this site.

For FTP and SFTP, enter the Host, Port, Username and Password (or a Private key for SFTP) your technical team gives you. A stored password or key is never shown again; leave the field blank to keep it.

The list shows each source's last run: how many files were found, ingested, unchanged, removed and failed, or Never run. To read a source right away, open its Actions menu and choose Read now.

Web crawlers​

A web crawler reads a public website, page by page, into a Web data source. Use it for your help center or product pages, so the assistant can answer from them and link to the page.

  1. Create a data source of type Web (see Create a data source).
  2. Open Artificial Intelligence > Web Crawlers and click Add Web Crawler.
  3. In Available Crawl Strategies, click Add on Sitemap.
  4. Fill in the fields below and click Save.
FieldWhat it does
NameThe name shown in the list.
Target Web data sourceThe Web data source that receives the pages. Several crawlers can feed one data source.
EnabledWhether the crawler runs on its schedule.
Re-index interval (minutes)How often the site is read again. Empty means once a day.
Base URLThe website to read, such as https://help.contoso.com. The crawler finds its sitemap.
Sitemap URL (optional)A sitemap address, when the crawler should start there.
Max pagesThe most pages read per run. Default 500.
Max concurrent requestsHow many pages are read at once. Default 4. Keep it low to be polite to the website.
Request timeout (seconds)How long to wait for one page. Default 30.
Include URL patterns / Exclude URL patternsWhich pages to read or skip, one pattern per line. Ask your technical team to write these. Exclude wins over include.
User-Agent (optional)How the crawler introduces itself to the website.

To read the site right away, open the crawler's Actions menu and choose Synchronize now. Only new, changed and removed pages are processed. Sources in answers link back to the original page.

How the assistant uses your knowledge​

  • Before each answer or on demand. With Enable preemptive retrieval-augmented generation (RAG) on (under Settings > Artificial Intelligence > Default Orchestrator), the assistant searches your knowledge before every answer. With it off, the assistant searches when it decides it needs to. Preemptive search is always on for profiles without tools.
  • Sources. Statements that come from your content get small numbers, and a numbered list of sources appears under the answer. Sources with a web address are links.
  • Staying on topic. Restrict answers to retrieved data only keeps the assistant to your content.
When answers are wrong or missing

Check that the data source has finished syncing, that the passage really contains the answer, and try a lower Strictness or more Retrieved documents. If the assistant still answers from general knowledge, turn on Restrict answers to retrieved data only and tell it in the System instructions to answer only from the provided content.