# Your document base, or why the agent answers beside the point

> An agent is only worth what it can read. The four defects that produce answers beside the point — duplicates, undated versions, scanned PDFs, implicit context — and the preparation that fixes them.

**Version HTML :** https://agence-intelligence-artificielle.eu/en/blog/why-your-agent-answers-beside-the-point  
**Langue :** en-GB

---

**On this page** The four defects What the work looks like What we then demand of the agent The « we’ll clean it up later » trap

Systems

# Your document base, or why the agent answers beside the point

When an agent answers beside the point, the cause is almost always the same: it was given documents nobody had reviewed in two years. The problem is not the model. It is the material.

2 September 20266 min readBy Sébastien Joumel

**The first job is rarely technical**In eight surveys out of ten, the work that unblocks everything is gathering, dating and de-duplicating documents that already existed.

## The four defects

### The duplicate

Three versions of the same procedure in three folders. The agent reads all three and answers with the one that most resembles the question — not with the right one.

### The undated version

A file called `procedure_v2_final_OK.docx` does not say whether it is current. With no date and no status, no machine — and no hurried human — can decide.

### The scanned PDF

An image of text is not text. Without optical character recognition those documents are invisible to the agent: it will not say it has not read them, it will answer without them.

### Implicit context

« Standard rates apply. » Standard for whom? What lives in people’s heads cannot be read. It is the most expensive defect, and the only one that requires writing.

## What the work looks like

It is neither long nor glamorous, and it is done with your teams rather than instead of them:

- inventory what exists and where — half a day, usually a surprise;
- keep one version of each document, dated, with a clear status;
- run the scans through OCR, and check the tables;
- write down what was nowhere: the list of documents per file type, the criteria, the known exceptions.

That last point pays off well beyond the agent: it is the same list you use to train a new joiner.

## What we then demand of the agent

Once the material is clean, the rule is simple: every answer cites the document and its version. Not a vague pointer — the exact reference, checkable in one click.

> An agent that cites its sources gets corrected. An agent that does not gets believed.

This has a useful side effect: it makes the gaps visible. When the agent answers « no current procedure covers this case », it has just flagged a real hole.

## The « we’ll clean it up later » trap

The temptation is to plug the agent into the documents as they are and tidy afterwards. It works for a week. Then the team hits three wrong answers, stops using the tool, and the project dies without anyone announcing it.

Better a smaller system on clean material than an ambitious one on a messy base. That is why the survey looks at the state of your documents before it talks about tools.

**Read on the site**

- [AI for manufacturing](https://agence-intelligence-artificielle.eu/en/ai-manufacturing) — ten years of quotes, never indexed
- [AI for legal teams](https://agence-intelligence-artificielle.eu/en/ai-legal) — the precedent bank to be dated
- [AI for construction](https://agence-intelligence-artificielle.eu/en/ai-construction) — yesterday’s bids feeding tomorrow’s
- [The survey of your ten systems](https://agence-intelligence-artificielle.eu/en/contact) — the state of your base is part of it

## Other notes

[Responsible AI · 14 August 2026« What if it makes things up? » — getting an agent to cite its sourcesA model produces a plausible sentence, not a true one. The four designs that make invention visible, and the single rule that matters: every claim points to its source.](https://agence-intelligence-artificielle.eu/en/blog/what-if-it-makes-things-up-getting-an-agent-to-cite)[Systems · 4 September 2026AI agent or plain automation: the £20,000 questionA third of the use cases brought to us need no model at all. How to tell a workflow from an agent, and why getting it wrong is expensive in both directions.](https://agence-intelligence-artificielle.eu/en/blog/ai-agent-or-plain-automation-the-20k-question)[Systems · 10 July 2026Getting cited by ChatGPT and Perplexity: what actually worksGenerative engines cite what they can read, structure and verify. Five concrete moves — including serving your pages as Markdown — and what is a waste of time.](https://agence-intelligence-artificielle.eu/en/blog/getting-cited-by-chatgpt-and-perplexity-what-works)

[All notes →](https://agence-intelligence-artificielle.eu/en/blog)

First call — fifteen minutes

## We will tell you which system to open first.

Describe the task that costs you the most. We will tell you what is feasible, how long it takes and what it costs. If the answer is no, you will leave with the two reasons why.

[Request the survey](https://agence-intelligence-artificielle.eu/en/contact) [Write to the agency](https://agence-intelligence-artificielle.eu/en/contact)

---

## Autres langues

- fr-FR : https://agence-intelligence-artificielle.eu/blog/socle-documentaire-pourquoi-votre-agent-repond-a-cote.md
- en-GB : https://agence-intelligence-artificielle.eu/en/blog/why-your-agent-answers-beside-the-point.md
- de-DE : https://agence-intelligence-artificielle.eu/de/blog/warum-ihr-agent-an-der-frage-vorbei-antwortet.md
- es-ES : https://agence-intelligence-artificielle.eu/es/blog/por-que-su-agente-responde-al-lado.md
- it-IT : https://agence-intelligence-artificielle.eu/it/blog/perche-il-vostro-agente-risponde-fuori-tema.md
- pt-PT : https://agence-intelligence-artificielle.eu/pt/blog/porque-o-seu-agente-responde-ao-lado.md
