The description of your documents already exists somewhere in your company. Your AI programme is rebuilding it
Gartner: through 2028, building your own unstructured metadata will cost 300% more than reusing what exists. The trade-off is settled at framing.
On 22 September 2026, at the Gartner Data & Analytics Summit in Mumbai, Prasad Pore put a number on something few AI programmes want to hear: through 2028, heads of AI, data science and data management will attempt to build their own unstructured metadata solutions, incurring costs more than 300% higher than they would if they used existing document and records solutions, skills and practices. The figure is striking. What actually decides the outcome is less so, and it is a matter of timing: this trade-off is settled during framing, and it does not reopen. Once the first in-house repository is in production, it becomes the reference. Nobody comes back, eighteen months later, to an asset teams have built, documented and presented to a steering committee.
This article is written for the CDO, CTO or head of knowledge management in a large group working to get its document estate ready for AI, whose programme has already run into its corpus. If your deployment is still a closed pilot on three hundred hand-picked documents, the argument does not bite yet: it will bite at the first scope extension. The person for whom this text is physically uncomfortable is not the reader. It is the AI programme director, the one who will be asked to reopen the “taxonomy and metadata” workstream that is already scoped, already staffed, already presented, and to admit that part of the work is being done twice. There is a precise moment where that discomfort is palpable, and most programmes have already lived it: the review where somebody asks who will maintain the repository once the project closes, and the answer is silence, followed by a commitment to revisit it in the next phase.
Two objections arrive before the end of the page. The first: “we are building our own repository, nobody knows our documents better than we do.” The second: “we already have the tool, our data catalogue is extending its scope to unstructured content, or our ECM has shipped an AI module.” Both are defensible. Neither answers the question the Gartner prediction raises, which is not about tooling but about duplication: where in the company is the description of this document already maintained, by whom, and as of when.
A word for regular readers of this blog. Recent pieces dealt with what two documents say to each other: which of them prevails when they disagree, and the scope within which a contradiction is even detectable. This one changes problem family. It is not about what your documents say, but about what your company has already written about them, and that your AI programme is now producing a second time.
One clarification before going further, because this article borrows its method from another discipline. Records management and archival practice have been writing down, for decades, what a document is, who answers for it, how long it lives and what supersedes it. What transfers here is the method, not the object: this is not about archiving, and this article recommends neither a records management system, nor an ECM, nor a corporate classification scheme. Nor is the records function the addressee, since it has already done this work. The addressee is whoever is funding a fresh description workstream today without having asked what already existed.
Before any tooling, one exercise fits in half a day and assumes no particular equipment. Take ten documents your assistant or agent actually served last month. If you have no usable access logs, take the ten documents your teams ask for most often; if you have neither, take the ten documents the business quotes in meetings, or nothing at all and start with a single document family. For each of them, put three questions to people who do not talk to each other: does a written description of this document already exist elsewhere in the company; who maintains it; how recent is it. Count how often the answer is “yes, and here it is.”
The reading rule fits in one sentence, and it has to be written down before you start, or the exercise yields nothing but a number. If the answer is “yes, and here it is” for most of the ten, your metadata workstream should be reopened and rebased on the existing owners before it is committed. If it holds for one or two, the workstream is legitimate, and the register then serves to record why, which beats never having asked.
What the Gartner figure says, and what it does not
The prediction is about comparative cost, not technical failure. It does not claim that in-house repositories collapse, or that data teams do poor work. It claims that at equivalent scope, rebuilding a descriptive apparatus over unstructured content costs more than reusing the one that exists. The same Mumbai session sets out two orders of magnitude that explain why the trade-off becomes visible now: by 2027, IT spending focused on multistructured data management would account for 40% of total spending on data management technologies and services, and from 2025 through 2029 the share of AI spending devoted to data readiness would increase sevenfold. The document estate is finally funded. The question is no longer whether there will be a budget, but which side of the line it falls on.
Some caution on this material: these are analyst predictions, describing a direction rather than measuring your case. They also say nothing about the method by which you would establish, in your own organisation, the gap between the two options.
The programme director deserves fairness here too. The fresh description workstream follows a delivery logic: it is the only part of the programme he can scope without depending on a business function’s calendar. A taxonomy is produced in sprints, with dated deliverables and a team he controls. Reusing what exists means going to ask people outside his reporting line what they hold, how, and since when. Between a workstream he controls and one he depends on, a programme under delivery pressure picks the first. The way programmes are cut up produces that choice mechanically.
One dated figure situates the moment, provided it is read for what it is. In its annual report published on 18 May 2026, drawing on a survey run in March 2026 among 1,000 decision makers in organisations of 1,000 employees or more across the US, UK, France and the DACH region, Nasuni reports that 16% of the organisations surveyed currently treat unstructured data management as a core IT investment, while 60% say they plan to invest over the next eighteen months. The source is a vendor of unstructured data platforms, so the figure counts above all as an admission about its own market, one it describes itself as largely unprioritised to date. And the window it points to is a season rather than an event: the budget framing cycles that open now and close on committed workstreams.
Getting your document estate ready for AI starts with an inventory nobody has done
In a group of more than five thousand people, document description rarely lives in one place, which is exactly why it is invisible. It lives in a business unit’s classification scheme, in the retention schedule held by the records function, in the mandatory metadata of a sector-specific ECM, in the document nomenclature of a certified quality process, in a legal department’s registers. None of these covers the whole estate. Together, they cover a share that nobody, in most organisations, has ever counted.
Let us name it once, and only once: description done twice. It is established by counting, on a real repository, how many documents already carry a description maintained elsewhere, by a named person, with an update date. The output is a count, and it is exactly the kind of measurement K-AI produces before any recommendation.
This is where a Document Knowledge Platform (DKP) comes in, in the precise sense of the triptych Govern (governing the document estate: ownership, authority, lifecycle), Clean (detecting and handling anomalies, duplicates, obsolescence and contradictions) and Activate (exposing the corpus to AI systems only once the first two are held). A word on categorisation, because the backdrop of this article belongs to an adjacent one: a DKP is not sold in the metadata management category and has no business appearing in the corresponding RFP. Metadata management is here the channel through which a document problem becomes visible and funded, nothing more. A DKP replaces neither a data catalogue, nor a data or AI governance platform, nor an ECM, nor an enterprise search engine: it runs before them, on the estate they consume. Its own unit of work, which neither an excellent records function nor an excellent data catalogue produces, fits in one line: measuring the state of the content itself across the living estate, reproducibly, and routing each case to a named expert who will decide. The records function describes and retains, the catalogue references and traces lineage, the DKP measures and routes. That position does not depend on any analyst’s publication calendar.
On the process itself, a word of reassurance, at the exact point where it is described. The material examined is document content, never conversation transcripts, usage logs or telemetry. The ingestion scope is contractual and bounded. Nothing is reused to train models. Scope is validated jointly by the business Document Owner and the CISO or DPO, never by IT alone.
What one real repository showed, and the boundary of what it establishes
At TotalEnergies Retail Power & Gas, whose experience report is public and named, roughly 500 pages of official web documentation fed the customer chatbot in production. Mapping and audit were carried out without migration: the pages stayed where they were. 19% of them needed correction, including cases that were very hard to spot by eye at that scale. 53% of the cases were resolved in three weeks, prioritising the most critical ones, with one to two business experts engaged half a day to a day per week. Thomas Bensoussan, Head of Digital Products, sums up the value in two effects: surfacing conflicts that are hard to spot by eye, and indicating which expert each case should go to. Business experts remained the decision makers throughout.
The boundary of this evidence has to be stated plainly, including where it weakens the argument. This experience report establishes prevalence on a specific repository at a specific moment: what was correctable, across those ~500 pages, during a first diagnostic. It establishes nothing about the question this article is concerned with: it does not document the comparative cost of an in-house metadata workstream, and it does not say whether a fresh descriptive asset was created or avoided, because that was not the mission’s object. What it contributes to the trade-off discussed here fits in one sentence: the business effort required to correct an existing repository was counted in half-days per week over three weeks, which gives an order of magnitude for what reuse costs. The other side of the comparison, what rebuilding costs in your organisation, only you can produce, and that is what the register below is for.
The descriptive reuse register
The half-day exercise produces one artefact, and only one. Call it the descriptive reuse register. Five columns: the document family, where a description already exists, the person or function who maintains it, the last known update date, and what the AI programme is about to recreate over that scope. A filled-in line looks like this:
General terms and conditions, business customer offer · Described in the legal department’s classification scheme and in the records function’s retention schedule · Maintained by the commercial legal domain lead · Last update declared at the previous pricing revision · The AI programme’s metadata workstream plans to rebuild its typing and versioning rules.
The register is not meant to become a standalone document: an artefact created for the occasion has neither owner nor lifecycle. Its lines live as an annex to the AI programme’s framing file, next to the AI use case register, and the framing lead maintains it, not a new function. The moment its first lines are written is decisive: never during the metadata sprint, that is, under the delivery pressure this article identifies as the cause of the problem. They are written cold, at framing, or at the next portfolio review, while the workstream is not yet committed. The half-day exercise and this register form a single exercise: the exercise produces the first lines, the register keeps them.
Two consequences to accept, because an exercise you cannot come out of bruised is not credible. If the register succeeds, meaning it shows that a description exists everywhere, well maintained and current, you have just documented that part of your metadata workstream is redundant, and somebody will have to stop it or rebase it in front of a committee that approved it. The cost is political, and it is real. Since the register is a dated written record, it also becomes a document that can be held against you later, should the rebasing turn out to be the wrong call. Writing it remains preferable, for one simple reason: the opposite decision, taken without a register, also leaves a trace, in the shape of a descriptive asset nobody can attach to an owner. This is a reasoned reading, with no established practice or published decision behind it; have your legal department qualify it before turning it into a standard.
On funding, this work does not require a new budget line: it attaches to the AI programme’s framing workstream, already approved, where data scope is decided. The counterpart applies immediately, because attaching creates a scheduling dependency: if framing slips, the register slips with it. Cut it by document family, so that each family completed is a deliverable in its own right.
And if you do nothing
The do-nothing option deserves to be stated, along with what it produces. The metadata workstream runs as planned. It delivers a fresh descriptive repository, coherent, properly documented, working for the programme’s use cases. The workstream will have met its commitments. What follows is quieter: that repository becomes a second descriptive apparatus, running alongside the ones that already existed, with its own update rhythm and its own owner. The two drift apart slowly. At the first quality control, the first vendor change or the first auditor question about the authoritative source, somebody has to say which of the two is the reference. And nobody has settled that question, because it belonged to the scope of neither workstream.
That responsibility has no holder today, and it does not call for creating a role. It is settled by a written attribution added to an existing one: process owner, domain lead, quality manager. A line added to a role description is enough. Worth noting in passing what K-AI answers for and what it does not: the counting, the reproducibility of the measurement and the routing to the right expert sit with the platform; the decision on which repository is authoritative stays with the business, and it will stay there.
Conclusion: audit, clean, monitor
The sequence does not change, and it starts earlier than most assume. Audit: count, on a real scope, what is already described elsewhere and by whom, before committing to any build. Clean: handle what is wrong, outdated or contradictory in the documents themselves, leaving the arbitration to business experts. Monitor: hold the measurement over time, because a document estate degrades again as soon as you stop looking at it. The programme director from the opening does not need to abandon his workstream. He needs to know, before committing it, how much of it redoes work that somebody in his company already maintains.
Frequently Asked Questions
Our data catalogue is extending its scope to unstructured content. Doesn’t that cover this?
A data catalogue describes assets and their lineage, and vendors in that category are indeed extending coverage to unstructured content. The boundary sits elsewhere: the catalogue references the document, it does not measure its state or say which of two competing documents is authoritative. That measurement happens upstream, on the content, and its result then feeds the catalogue.
Why not simply hand this to our records function?
Because it is a different request. The records function holds retention and regulatory description over a defined scope; it has neither the mandate nor the means to qualify, across the living estate, what feeds an AI assistant. It does, however, hold part of the answer to the register’s first question, and it is usually the function the AI programme never called.
What is the confidentiality framework for such a review?
The ingestion scope is contractual and bounded to designated document content, excluding transcripts, usage logs and telemetry. No customer data is reused to train models. Scope is validated jointly by the business Document Owner and the CISO or DPO, never by IT alone.
How long before something usable appears?
That depends on the scope chosen and on business expert availability, two parameters that belong to you. The available public reference point is the TotalEnergies Retail Power & Gas repository: across roughly 500 web pages, 53% of the identified cases were resolved in three weeks with one to two experts engaged half a day to a day per week. It is a dated order of magnitude on a known scope, and it does not transfer as is.
Do we have to migrate our documents to run this exercise?
No. The TotalEnergies Retail Power & Gas audit was carried out without migration, the pages staying in place. That is precisely what separates this from a document overhaul project: it moves nothing, it measures and routes.
Sources
- Gartner Data & Analytics Summit 2026 India: Day 2 Highlights — Gartner, Mumbai, 22 September 2026 (Prasad Pore session; predictions on unstructured metadata, multistructured data spending and AI data readiness).
- Nasuni Research Finds 97% of Enterprises Are Adopting AI Agents, Yet Most Projects Fail to Meet Objectives — Nasuni, 18 May 2026; Sapio Research survey run in March 2026 among 1,000 decision makers in organisations of 1,000+ employees (US, UK, France, DACH).
- K-AI — customers page and TotalEnergies Retail Power & Gas experience report — source of the RPG repository figures cited in this article.
Where to Go From Here
K-AI Corpus Diagnostic — 10 business days on your document estate, full report of the 20 most critical anomalies, money-back guarantee if no meaningful anomaly is found. A one-hour conversation fills in the first lines of your descriptive reuse register: on a document family of your choice, what is already described elsewhere, who maintains it, and what your AI programme is about to rebuild on top of it. Reach the K-AI team: contact@k-ai.ai. The scope of every diagnostic is validated jointly by the business Document Owner and the CISO/DPO, never by IT alone.
K-AI already works with CMA CGM, Veolia, PwC, BNP Paribas, TotalEnergies and CEVA Logistics. Partners: AWS, Snowflake, Microsoft, Wavestone, Devoteam.
