← All news
Press · September 7, 2026 · 11 min read

Authoritative sources for enterprise AI: what you certify when you designate one

Authoritative sources for enterprise AI: what you certify when you designate one

SharePoint now lets you mark a site authoritative for Copilot. The signal sits on the container; the contradiction lives in the documents.

Since mid-June 2026, preparing a Microsoft 365 environment for an internal assistant includes a step that did not exist last year: designating selected SharePoint sites as authoritative, so that Copilot gives priority to their content. Administrative documentation describes the setting as a signal telling search engines and AI agents that the site’s content is official, trusted and verified.

Someone in your organisation will have to set that signal. That person holds a job title, and has not read the thousands of documents the site contains.

Two instincts follow, and neither is sufficient. The first is to lean on the in-house document register: the list of “official” sites maintained by the information management function, which records who publishes where and never whether two publications contradict each other. The second is to assume the tool already purchased covers this, a data catalogue, a governance platform or a DMS, none of which reaches down to the sentence that governs. Our 26 August article dealt with how document readiness is assessed; this one deals with a single operational act that puts a named person behind a verifiable claim.

Our thesis: designating an authoritative source does not reduce the ambiguity in your document estate, it authenticates it. And the only defensible basis for that designation is a prior count.

An AI readiness prerequisite landed in administrators’ hands

The fact is dated and documented. Message centre notification MC1310687, dated 14 May 2026, introduces authoritative sites for SharePoint Online in Microsoft 365 Copilot; rollout completes for targeted release tenants by the end of May 2026, with general availability following in mid to late June 2026.

Three details matter before any interpretation. The setting is applied site by site, through the IsAuthoritative property of the Set-SPOSite administrative command, with a graphical admin surface announced to follow. It is off by default: no tenant has an authoritative site until someone decides on one. And it requires a premium Copilot licence. So this is not a box already ticked that ought to be unticked. It is a decision, taken at a dated moment, by an identifiable person.

The capability deserves to be taken seriously, in its strongest form. It answers a real and long-standing problem: until now, nothing allowed an organisation to tell an internal assistant that the copy of a file sitting on the HR portal outranks the one lingering in a project workspace. It is the positive counterpart to the symmetrical control, Restricted Content Discovery, which instead removes selected sites from search and discovery. It encodes in the tooling a notion information managers have worked with for decades, the source of record. And the ecosystem is not overselling it: the deployment doctrine published on 4 September 2026 by Tony Redmond, an independent technical reference, states the problem plainly.

“AI cannot distinguish between authoritative information and outdated, inaccurate, or duplicated content. Files are files, and if the information in Microsoft Search and the semantic index leads Copilot to a file, it will use it to generate responses to user prompts.”

The same piece lists four remedies. Three address access and sensitivity: removing sites from discovery, archiving, applying loss-prevention policies to sensitivity labels and prompts. The fourth, designating authoritative sites, is the only one that addresses veracity. It is also the only one that cannot be verified, because it consists of asserting.

The signal sits at site level; the contradiction lives at document level

This is the granularity gap that decides everything, and it is not vendor-specific. Any trusted-source designation applied at container scale, whether that container is a site, a workspace, a connector or a collection, inherits the same limit: it extends a presumption of authority to everything the container holds. Microsoft is the first dated, documented instance at office-suite scale, not the special case.

In practice, marking a site authoritative extends that presumption to the two divergent versions of the same procedure it hosts. The signal attests that both come from a place that governs, without identifying which one governs. The ambiguity is not resolved, it is certified, and it now surfaces at the top of the list.

At the top of the list, and flagged as such. This is what a purely technical reading misses: results from an authoritative site appear first, carrying a trust marker indicating they come from the organisation. The certification does not stay in the admin console; it appears on the screen of every employee who asks a question. A result placed first and marked as official is consumed as an answer, not as a lead. That transfer is what makes the signature heavy, more than the granularity argument: whoever sets the signal commits the confidence of thousands of readers who will not check.

The limit is a property of the mechanism rather than an implementation defect. It has a direct bearing on how the feature should be read: documentation states that this class of setting concerns ranking and trust, and does not constitute a security boundary. Our 10 July article addressed exactly that confusion between governance and access control. The act discussed here is not one more access control. Access says who may read; it does not say which of the two readable documents is the right one. Two procedures equally authorised, equally non-sensitive, equally accessible and mutually contradictory pass untouched through all four remedies.

That leaves the option of designating nothing, which is the status quo in every environment. It protects nothing: the assistant keeps using whatever files it finds, with no hierarchy and no trust marker displayed. The choice is therefore not between certifying and abstaining. It is between certifying at random and certifying after a count.

What a Document Knowledge Platform produces, and what the designation assumes

A Document Knowledge Platform (DKP) is the document quality and governance layer that sits upstream of AI systems. It holds three functions: Govern, establishing ownership, authority and lifecycle across the estate; Clean, detecting and resolving anomalies, divergent duplicates, obsolescence and contradictions; Activate, opening the corpus to AI systems only once the first two hold. This position does not depend on vendor release calendars or analyst publication cycles: the document layer precedes the engine, whatever setting ships this quarter or whichever quadrant is published next. Gartner, for its part, names unstructured data governance as an AI readiness workstream in its strategic roadmap published in August 2025.

A DKP does not replace a data catalogue, a data and AI governance platform, a DMS or an enterprise search engine. It runs before them, on the estate they consume, and produces the factual object missing at the moment of designation: a count of the documents claiming authority on the same questions, within the scope under consideration.

The process covers document content only: no conversation transcripts, no usage logs, no telemetry. The ingestion scope is contractual and bounded to designated sources, and the corpus is never reused to train models. Business experts remain the decision-makers: the diagnostic establishes the cases and routes them to the right expert, who then decides.

What a count produces, in practice. On the reference set behind the TotalEnergies Retail Power & Gas customer chatbot, roughly 500 pages of official web documentation were mapped and audited without migration: 19% of pages required correction, including cases that are hard to spot by eye at that scale. 53% of cases were resolved within three weeks, prioritising the most critical ones, with one to two business experts contributing half a day to a day per week, and chatbot accuracy measured before and after. On an anonymised case in energy and industry, an initial diagnostic detected 398 document conflicts. These figures are counts taken over a bounded scope, where market surveys on data readiness measure declarations and return ranges that vary widely with the scope chosen.

Before designating a site as authoritative: take the twenty questions most often put to your internal assistant, and count, within that one site, how many documents claim to answer them authoritatively. If the number exceeds one for several of those questions, the signal is about to certify a contradiction.

What still applies, and who answers for the designation

The European regulatory calendar has been eased without disappearing. Regulation (EU) 2026/1744, known as the Digital Omnibus, published on 24 July 2026 and in force since 27 July 2026, postpones obligations for stand-alone high-risk systems under Annex III to 2 December 2027, and those under Annex I to 2 August 2028. What still applies today: supervision and enforcement by national authorities began in early August 2026; the AI literacy obligation in Article 4 has applied to providers and deployers since February 2025, in its wording as amended in July 2026; the transparency obligations in Article 50 have applied since 2 August 2026. The postponement moves an enforcement date; the substance of the requirements is unchanged.

That said, the designation question is not primarily regulatory. An internal assistant at a large non-tech group most often falls outside Annex III. The signal, however, creates a claim of officiality that lives inside the organisation, independent of any legal classification, and that will have to be owned in front of the business functions whose documents were promoted. Three roles therefore need separating and naming, following the mapping we apply with our clients: the Document Owner in the business, who answers for the content; the Authority, who holds the mandate to designate a source of record; the Steward, who maintains the site over time. Without those three roles, “authoritative” is a label with no one accountable. It is also why a diagnostic engagement is approved jointly by the business Document Owner and the CISO or DPO, never by IT alone.

Conclusion

The step is a good one, and it should be taken. But in the right order. Audit the scope actually exposed, to know how many documents in it claim to govern. Clean the contradictions, divergent duplicates and out-of-date content the count surfaces, leaving the judgement calls to business experts. Monitor thereafter, because a designated site drifts from the first unreviewed publication onward. Designating before auditing means putting an official stamp on an ambiguity, then displaying it to the entire company.

Frequently Asked Questions

What is an authoritative source for enterprise AI?

A site designated as a source of record, so that enterprise search engines and internal assistants give priority to its content in their answers. On SharePoint Online, the designation is applied site by site and has been available since mid-2026; it is off by default. It is a trust and ranking signal, not a security measure and not a verification of content.

Is designating a site authoritative enough to make an internal assistant reliable?

No. The signal applies to the container, while contradictions live in the documents. If a site holds two divergent versions of the same procedure, designating it pushes both to the top of the list with a trust marker shown to the user. Reliability requires work at document level.

Who in the organisation should decide that a site is authoritative?

The mandate belongs to the business function that owns the content, not to the team holding technical access to the setting. Three roles usefully separate: the Document Owner who answers for the content, the Authority who holds the designation mandate, the Steward who maintains the site. The technical setting merely executes the decision.

How do you know what a site actually contains before designating it?

By counting it. A corpus audit establishes, over a bounded scope, how many documents address the same questions, which of them contradict each other, which are divergent duplicates and which are out of date. The result is a counted, defensible record, where an opinion remains arguable.

Does the diagnostic require our documents to leave our environment?

The ingestion scope is contractual and limited to designated sources. It covers document content only, excluding transcripts, usage logs and telemetry, and the corpus is never reused to train models. Approval is joint, by the business Document Owner and the CISO or DPO, never by IT alone.

Sources


Where to Go From Here

K-AI Corpus Diagnostic — 10 business days on your document estate, full report of the 20 most critical anomalies within the scope actually exposed to your AI systems, money-back guarantee if no meaningful anomaly is found. To count what a site holds before designating it a source of record, reach the K-AI team: contact@k-ai.ai. The scope of every diagnostic is validated jointly by the business Document Owner and the CISO/DPO, never by IT alone.

K-AI already works with CMA CGM, Veolia, PwC, BNP Paribas, TotalEnergies and CEVA Logistics. Partners: AWS, Snowflake, Microsoft, Wavestone, Devoteam.

And in your organization, what does your document estate look like?

30 minutes with a founder. We audit a sample of your documents for free and show you exactly what K-AI detects.

Book a demo → Read other articles