Is Your Knowledge Base Ready for AI? A 12-Point Content Audit
Score your knowledge base before connecting it to an AI assistant: audit authority, freshness, structure, metadata, permissions, testing, escalation, and governance.

An AI assistant can retrieve information in seconds. That does not mean the information deserves to become an answer.
Many knowledge bases were built for people who can interpret an old date, notice that two policies disagree, or ask a colleague what a vague sentence really means. An AI assistant has no such organisational instinct. If you connect it to every page, PDF, slide deck, and shared folder without preparation, you do not create one reliable source of truth. You make every hidden content problem available at conversational speed.
This audit helps you decide whether your content is ready for a website AI assistant, an internal copilot, or another knowledge-grounded experience. It deliberately evaluates the information before the model. A sophisticated retrieval system can find content efficiently; it cannot decide which of two conflicting policies your company intended to honour.
First: what “AI-ready” actually means
AI-ready content is not content rewritten to sound as if a machine produced it. It is information that is:
- authoritative enough to support an answer;
- clear about where, when, and to whom it applies;
- structured so the relevant passage can be found without losing its conditions;
- accessible only to the people and systems permitted to use it;
- owned, tested, and maintained as the business changes.
Many assistants use retrieval-augmented generation, usually called RAG. In plain language, the system searches an approved collection when a question arrives, retrieves relevant passages, and gives those passages to a language model to compose a response. This is useful because answers can draw on your current business information without retraining the model for every change.
But RAG is not a truth machine. Retrieval can surface an outdated page, detach an exception from a table, or find content meant for the wrong market. The quality of the answer is limited by the quality and control of what can be retrieved.
How to score the audit
Score each of the 12 checks from 0 to 2:
- 0 — Missing: there is no dependable process or evidence.
- 1 — Partial: the practice exists for some important content but is inconsistent.
- 2 — Controlled: the practice is documented, applied to the intended scope, and can be demonstrated.
Do not score from memory. Sample the actual sources planned for launch and record evidence: URLs, document IDs, owners, review dates, test results, or access rules. When different repositories vary widely, score the weakest one that the assistant will be allowed to use.
1. Is there an authority order for sources?
An assistant needs to know more than which documents are searchable. It needs a rule for which source wins.
Create a source register that marks each item as canonical, supporting, historical, or excluded. Define an authority order by topic: for example, the approved returns policy may outrank a help article, while the product catalogue may outrank a campaign landing page for technical specifications.
Then test conflicts deliberately. If two approved sources disagree, the system should not quietly choose the passage with the closest wording. Resolve the conflict in the content or define a safe behaviour such as stating that the answer cannot be verified and escalating to the policy owner.
Score 2 when: canonical sources and conflict rules are documented by topic, and known contradictions have an owner.
2. Is the assistant's scope explicit?
“All company knowledge” is not a useful launch boundary. List the audiences, intents, products, markets, channels, and actions the assistant is expected to support. List exclusions just as clearly.
A public sales assistant might cover product fit, pricing principles, integrations, and booking, while excluding account-specific support, contractual interpretation, and unpublished roadmap questions. Those exclusions determine which repositories should never enter the index and which questions should trigger a handoff.
Scope also prevents accidental answers across contexts. A partner discount, an internal workaround, and a public offer may all be correct inside their own boundaries and harmful outside them.
Score 2 when: every included source supports a named use case and audience, and out-of-scope topics have defined responses.
3. Does every important source have an owner and review date?
“Last modified” is not the same as “last verified.” A formatting edit yesterday does not prove that a policy is current.
For each high-value source, record a business owner, last substantive review, next review or expiry date, and status. Review frequency should reflect the subject: rapidly changing prices may need a different cycle from stable company history.
Define what happens when content expires. Options include removing it from retrieval, lowering its authority, showing a warning to reviewers, or blocking answers until the owner confirms it. A review reminder with no consequence is not governance.
Score 2 when: critical content has accountable owners and review deadlines, and overdue material is automatically surfaced or excluded.
4. Can the content answer a real question directly?
Content written for navigation, branding, or legal completeness may mention a topic without answering what a visitor asks.
Take a sample of real questions and locate the exact passage that should support each answer. Check whether that passage contains the decision, key conditions, exceptions, and next step. If the answer requires a person to combine five pages and infer an unstated rule, it is not ready simply because all five pages are indexed.
Prefer clear headings, descriptive labels, complete sentences, and one primary purpose per section. Define acronyms and internal terms. Put the important qualification beside the claim it limits; do not hide it in a distant footnote.
Score 2 when: common questions map to self-contained, unambiguous passages that include their necessary conditions.
Related: How to Launch a Website AI Assistant in 30 Days Without Losing Customer Trust
5. Will structure survive extraction and chunking?
Knowledge systems often split documents into smaller passages for retrieval. If a heading, row label, or exception is separated from its content, the retrieved fragment may become misleading.
Inspect PDFs, tables, scanned pages, slide decks, screenshots, accordions, and videos. Confirm that meaningful text is actually extractable and that reading order is preserved. Give tables descriptive headers, repeat essential context where necessary, and provide text equivalents for information trapped in diagrams or screenshots.
Avoid designing around an arbitrary “ideal chunk size” before testing your material. The important question is whether each retrieved unit remains semantically complete: does “available for 30 days” still identify what is available, to whom, and from which event the 30 days are counted?
Score 2 when: representative files extract cleanly, and headings, lists, tables, notes, and exceptions keep their meaning in retrieved passages.
6. Does metadata identify the correct context?
Similarity alone is not enough. Two pages can use nearly identical language while applying to different countries, customer tiers, products, or dates.
Attach metadata that allows retrieval to filter or rank content appropriately. Useful fields often include:
- product or service;
- audience or account type;
- country, market, or jurisdiction;
- language and locale;
- effective date, expiry date, and version;
- content owner and approval status;
- confidentiality or access class;
- source type and authority level.
Use controlled values rather than improvised labels such as “UK,” “United Kingdom,” and “GB” for the same market. Missing metadata should fail visibly; it should not silently become “applies everywhere.”
Score 2 when: required metadata has a defined schema, valid values, quality checks, and a role in retrieval decisions.
7. Are privacy and access permissions preserved?
Being technically reachable does not make content appropriate for every answer. Before ingestion, classify sources as public, internal, confidential, personal, or otherwise restricted according to your organisation's rules.
The assistant should enforce access at retrieval time, not merely hide sensitive text after it has been retrieved. Verify user identity where needed, preserve source permissions, separate public and internal collections when appropriate, and use the minimum data required for the use case.
Also review conversation data: what is stored, who can review it, how long it is retained, and whether personal information is needed at all. Remove secrets, credentials, unneeded personal data, private comments, and tracked changes from source files before indexing.
Score 2 when: source access is mapped to user permissions, sensitive data has been reviewed, and both source and conversation retention are controlled.
8. Are language and market versions genuinely aligned?
Translation is not only a writing task. A Turkish page, Arabic PDF, and English policy may have different update dates or legal scope even when their titles look equivalent.
Link translations to one canonical content record. Mark whether each version is a direct translation, a local adaptation, or a separate market policy. Preserve product names and approved terminology, but do not assume that one language can safely fill a gap in another.
Test cross-language questions, mixed-language product terms, local date and number formats, and right-to-left rendering where relevant. Native review matters most for conditions, exclusions, and calls to action—not just the opening paragraph.
Score 2 when: language versions are connected, their equivalence or differences are explicit, and updates cannot leave silent gaps.
9. Can an answer be traced back to its source?
Your team needs to investigate why an answer appeared. Visitors may also benefit from a link when they want the full policy, technical detail, or original context.
Keep stable source identifiers, useful titles, section anchors, version information, and canonical URLs through the ingestion process. Decide when the experience should display sources and how employees can inspect the passages retrieved for any answer.
Traceability is not a substitute for correctness: a citation to the wrong policy is still wrong. It is the evidence layer that makes review, correction, and trust possible.
Score 2 when: reviewers can move from an answer to the exact source passage and its current owner, version, and status.
Related: What Does a Website AI Assistant Cost? The Total-Cost Worksheet
10. Do you have a test set made from real questions?
A few polished demo prompts will not reveal whether the knowledge base works for visitors.
Build a test set from anonymized site searches, support tickets, sales calls, contact forms, and input from customer-facing teams. Include common wording, misspellings, vague questions, multi-part requests, local terminology, and questions containing false assumptions. Add near-neighbours: questions that sound similar but should produce different answers because the market, product, date, or account type changes.
For every test, record the expected answer elements, acceptable sources, required clarification, prohibited claims, and correct handoff. Run the set before launch and after changes. Score answer correctness and completeness separately from retrieval quality so that you know whether to repair the content, search setup, or response behaviour.
Score 2 when: a representative test set has expected outcomes, owners, pass criteria, and repeatable regression runs.
11. Are gaps and escalation paths designed?
No knowledge base will cover every question. Readiness means handling absence safely, not pretending absence can be eliminated.
Define the conditions for “I cannot verify that,” a clarifying question, a link to a process, or a human handoff. Give the receiving team the visitor's question, useful conversation context, sources already checked, and the reason for escalation—subject to appropriate consent and data controls.
Record unanswered and low-confidence questions as a content backlog. Separate genuine content gaps from questions that are intentionally out of scope. Otherwise teams may create risky content simply to improve an automation rate.
Score 2 when: uncertainty produces a useful next step, handoffs preserve context, and knowledge gaps enter an owned improvement process.
12. Is change governed after launch?
The knowledge base will change; the question is whether the assistant changes with it safely.
Document the route from content draft to approval, publication, ingestion, testing, and rollback. For material changes, identify which test questions are affected and who signs off. Monitor sync failures and orphaned sources. Keep an audit trail for additions, deletions, permission changes, and authority changes.
Use conversation reviews to improve source content, not only the prompt. A recurring incorrect answer may point to duplicate pages, a weak heading, missing metadata, or an unresolved business rule. Fix the cause, rerun the relevant tests, then confirm that the updated source is actually live.
Score 2 when: content and index changes follow an owned release process with monitoring, regression tests, and a workable rollback.
The 24-point scorecard
| Audit area | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 1. Source authority | No hierarchy | Informal or incomplete | Documented and tested | |
| 2. Scope | Open-ended | Broad boundaries | Named uses and exclusions | |
| 3. Owners and freshness | Unknown | Partial coverage | Owners, reviews, expiry rules | |
| 4. Answerability | Information is implied | Some direct answers | Clear, complete answer passages | |
| 5. Structure and extraction | Untested | Main formats mostly work | Representative formats verified | |
| 6. Metadata | Missing | Inconsistent | Controlled schema used in retrieval | |
| 7. Privacy and access | Not mapped | Manual or partial controls | Permissions enforced and tested | |
| 8. Language and market | Unlinked copies | Some alignment | Connected, governed variants | |
| 9. Traceability | Source unknown | Document-level only | Exact passage, version, and owner | |
| 10. Real-question testing | Demo prompts only | Small ad hoc set | Representative regression suite | |
| 11. Gaps and escalation | Dead ends or guesses | Generic handoff | Contextual handoff and owned backlog | |
| 12. Change governance | Unmanaged sync | Basic publishing routine | Approval, monitoring, tests, rollback |
Add the 12 scores, then interpret the total:
- 0–8: Stop and repair the foundation. A live assistant would mostly expose unresolved content and access problems. Choose one use case and build a clean canonical collection first.
- 9–16: Ready for a constrained internal test. Limit the audience and sources. Use the pilot to resolve specific gaps, not to claim broad readiness.
- 17–21: Ready for a controlled external pilot. Launch a narrow use case with monitored conversations, clear handoffs, and release gates for the remaining weak areas.
- 22–24: Strong content readiness. Proceed with technical and safety testing, but do not treat the score as a guarantee. Retrieval, model behaviour, integrations, and the complete customer journey still need evaluation.
Regardless of the total, treat three findings as launch blockers: the assistant can retrieve content a user is not allowed to see; critical sources contradict one another without a resolution rule; or high-impact content has no accountable owner.
A small before-and-after example
Imagine this sentence buried in a general shipping page:
Before: Express delivery is usually available, although some orders and locations may take longer. Contact us for details.
It sounds reasonable but leaves an assistant to guess. Which locations? What does “usually” mean? Which orders are excluded? Is this a promise or an estimate?
A retrieval-ready version makes the decision and its boundaries explicit:
After — Express delivery eligibility (Türkiye, online retail orders): Express delivery can be selected at checkout when the destination postcode and every item in the basket are eligible. It is not available for made-to-order items or deliveries outside the listed service area. The delivery estimate shown at checkout is the current estimate for that order; do not promise an earlier date. If no express option appears, offer standard delivery or transfer the customer to the delivery team.
The improved version is not “written for robots.” It is better operational content for people too. It names the market and order type, states the decision point, keeps exclusions beside the claim, identifies the authoritative live value, prevents an unsupported promise, and supplies a fallback.
Turn the score into a repair plan
Do not begin by rewriting the entire company wiki. Start with the intended customer journey and the questions that matter inside it.
- Select the first use case and remove unrelated sources from scope.
- Fix all zero-score items that affect permissions, authority, or high-impact answers.
- Convert the most frequent real questions into the initial test set.
- Improve the smallest group of sources needed to pass those tests.
- Run retrieval and answer evaluation, then review failures by cause.
- Launch narrowly, observe real conversations, and add content only through the governed process.
The outcome of the audit is not a prettier document library. It is a dependable chain from a visitor's question to the right approved passage, a bounded answer, and a safe next step.
Orbitra helps organisations prepare the knowledge, multilingual rules, evaluations, and human handoffs behind a website AI assistant. If your audit exposes weak ownership, conflicting sources, or unclear scope, we can help turn one high-value journey into a controlled first release—without pretending that indexing every file is the same as being ready.