Now the uncomfortable part: you can't just throw fifty years of accumulated files at an AI and expect good answers. Garbage in equals garbage out, and the companies disappointed by AI are mostly companies that skipped this step. The good news is that readiness is a process with a known shape, and it's smaller than it sounds.
Why can't AI just sort out messy data itself?
Because AI retrieves and reasons over what exists; it can't tell which of your three conflicting policy documents is the real one. If the shared drive holds two versions of a spec, an obsolete procedure with no marking, and a folder structure only its creator understood, the system will faithfully surface conflicting answers. The failure won't look like an error; it'll look like an answer someone trusts and shouldn't.
That's why readiness, not model quality, is what separates a working deployment from a disappointing one. The AI is the easy part. The hard part is deciding what's true.
What does "AI-ready" actually require?
Less than a cleanup of everything: it requires owned, current sources for the domains that matter. The working model from the manufacturer above: appoint knowledge masters, one or two per department, each accountable for their domain's content: which documents are authoritative, which are obsolete, what's missing and needs writing down. About ten people across the company, doing it as part of their job rather than as a task force.
Pair that with a file structure people actually use. A sensible shared-drive hierarchy, agreed conventions for naming and versioning, and a habit: when a document is touched, it lands in the right place, marked current. Files migrate into order naturally as they're used, instead of through a heroic one-time scrub of fifty years of history, most of which nobody will ever query anyway.
Do you have to consolidate four ERPs first?
No, and this is the misconception that stalls the most projects. The systems stay where they are. An AI layer connects to all four eras of ERP, the CAD metadata, the CRM, and the file shares, and unifies them at the answer level rather than the database level. Migration projects that consolidate decades of records into one system are exactly the six-figure, multi-year efforts a layer exists to avoid.
What the layer can't do is adjudicate truth, which is why the knowledge-master work and the layer go together. The people decide what's authoritative; the system makes the authoritative version findable by everyone.
What's a realistic sequence?
Start narrow, prove it, widen. Pick one domain with an owner and contained sources: HR policies are a classic, engineering standards another. Have the knowledge master mark what's current, retire the duplicates, fill the worst gaps. Connect it, and let the department feel what a trustworthy answer is like. Then repeat, domain by domain, with the file-structure habits running underneath.
Within a few cycles something useful happens culturally: people stop treating the shared drive as a junk drawer, because it's become the thing that answers their questions. Readiness stops being a project and becomes maintenance.
How do you know when you're ready enough?
When the domains people query most have a single, owned, current source; that's the bar, not perfection. You'll never have fifty years of files in perfect order, and you don't need to. You need the questions people actually ask to land on answers someone is accountable for. Everything else can stay sediment.
