On 17 September we ran a webinar for chief geologists, technical directors, exploration leads and resource modellers. The question in the title was blunter than usual: you already digitised, so why can you still not ask it anything? Vlad Tiazhov took the room through the engine room, and Chris Yeomans of Neo Geo Consulting came in as an independent geologist to argue with both of us. The full recording is below, and the technical white paper that sits behind it is at the foot of this page.

Fifteen minutes in, Chris said the most useful sentence anyone has said to me about AI in geology. Every answer you get from an AI is a hallucination. It is just that the ones you do not think are hallucinations happen to fit your worldview.

That is not a complaint about the technology. It is a description of what the technology is. A large language model is a probabilistic engine. It predicts the next token. Point one at a folder of raw geological PDFs and it will produce a confident, well-formatted answer whether or not anything underneath it is real, because agreement is the likely next token. Ask a general model for every ASX company under half a billion dollars of market capitalisation ranked by enterprise value per ounce and it will hand back five or six names when the real answer is closer to 120. It is not lying. It is doing exactly what it was built to do, on data that was never put in a condition for it to work on.

What breaks a geological record?

Recency, first. If a figure is used dynamically it has to be updateable, and a right number from the wrong vintage is a wrong number.

Then units of measure. Grade is reported in grams per tonne, in per cent, in parts per million, and in the volumetric conventions that a particular placer assemblage in a particular country attracted in the 1970s. Currency and economic convention behave the same way. Then reporting standards, because JORC, NI 43-101, SAMREC, S-K 1300 and the historical Soviet GKZ are not interchangeable and cannot be mapped to each other by lookup table.

Then the geospatial layer, which thins out the older the record gets. A coordinate referenced to the Pulkovo Meridian, the geographical reference of the Russian Empire from 1844, sitting 30 degrees 19 minutes east of Greenwich, is not a coordinate a modern GIS can place. Then a century and a half of evolving measurement convention, mathematics and interpretation. And then the problem that sits on top of all of them: credible sources that disagree with each other, which forces the system to hold a view on which one to trust rather than averaging them into something that is true of nothing.

Each one is theoretically solvable. None of them is solvable once and then forgotten about.

Why does extracting the data accurately not solve it?

This is the part that surprises people, and it is worth sitting with.

Soviet-era drill logs are recorded metre by metre. Assume the extraction is perfect, 100 per cent, every character correct. Now ask for every intercept over five metres grading better than two per cent copper across a corpus of millions of those rows. The model has to compute its way to that answer, row by row, because the question does not match the shape the data is stored in. On a corpus of any size it will not finish, and it will not be reliable if it does.

Normalised to modern interval-based reporting, the same question is a query. That is the entire difference between a digitised archive and a usable one, and it is why "we already digitised" and "we can ask it things" are two completely different states.

How accurate does extraction have to be?

Requests arriving with a pilot brief average around 90 per cent on extraction completeness and accuracy. That threshold is unusable at scale, and the reason is structural rather than a matter of taste. Errors at the 90 per cent level cannot be filtered, because nothing distinguishes the correct rows from the incorrect ones without re-verifying every cell. A team running on 90 per cent has done a fast pull and inherited the burden of cell-by-cell verification on every row that matters. The work is still in front of them.

Mining-grade extraction has to push past 99 per cent on the metrics that drive decisions, on real corpora rather than on the cleanly printed examples chosen for vendor benchmarks, and it has to know which one per cent it could not process and flag that for human review.

A tier-one producer commissioned a pilot on complex Soviet-era geological reports from their own historic exploration archive: Cyrillic, multi-format, historical. Seventy-one Priority-1 documents, 164 input files, the hardest single drill-log passports carrying up to 20,000 data points apiece. The brief was to meet or exceed every threshold in the producer's own scoring framework, on real data with real ground truth.

MetricThresholdResult
Depth interval accuracy≥ 90%99.65% row recall
Assay value accuracy≥ 90%99.21% combined element-cell accuracy
Precision≥ 90%99.06% aggregate
Recall≥ 90%99.65%
Character recognition, word≥ 85%~97% word, ~99% character
Lithology pattern recognition≥ 80%96.3% of assay rows enriched
Processing failure rate< 5%0 of 164 hard failures

Scored ground-truth result against the producer's own thresholds. Anonymised tier-one producer pilot, 2026.

No single technique reaches those numbers. The result came from combining computer vision, classical character recognition, model-based extraction and statistical reconciliation, validated through a benchmarking framework with fifty to two hundred hard samples per pipeline step and regression-tested on every change. That benchmarking discipline, rather than the model selection, is the part most internal builds skip.

An operating geologist on a frontier project put the honest version of this better than we could: at this point I am not demanding 99 per cent accuracy, I want to know whether it is 70, 80 or 90, and then we work in that context. Confidence in the number matters more than the number. A vendor quoting 99 per cent without naming the corpus, the metric and the validation method is selling marketing.

How do you georeference a hand-drawn map from 1925?

A map with no coordinates is not a map a GIS can use, and most historical exploration archives are full of them. The method runs in two passes.

The first digitises the map itself and reads the settlements off it, then matches those names to their modern equivalents, allowing for the towns that were renamed. That gets close and leaves three to five per cent drift, which at map scale can be a different village.

The second pass uses the railways. Almost every hand-drawn Soviet sheet carries the rail line, and rail geometry is distinctive, slow to change and available as a global layer. Align the curve of the line on the sheet to the curve of the line in the modern layer and the sheet lands. On a first test set, 30 of 30 maps were placed in the correct location, at about two per cent drift against a manually created ground truth. The method is agnostic. It travels to any old map, anywhere, and it is close to what a Kazakh geologist does by hand, done at the scale of an archive.

How does Pulse decide which source to trust?

An assay certificate, a company announcement, a historical interpretation, a corporate presentation and management commentary are not equal evidence, and a system that treats them as equal will produce a number that cannot be defended in a technical review.

Every document type is categorised separately at extraction, by report type, and then ranked for credibility with technical reports at the top. Figures are cross-referenced against other credible sources, and it is that friction between sources, rather than any single extraction, that produces a number worth putting in a model. Every value is timestamped and opens the source page it came from in one click. That validated layer is what we call AuthentiQ.

There are no unofficial sources in the platform. Where we track what is said about an asset beyond company disclosure, that is separate risk work for royalty clients and it is never blended into a disclosed figure.

What does this cost to run?

The reassuring story is that AI gets cheaper every year. Independent analysis from Epoch AI puts the fall in per-token price at roughly an order of magnitude a year, and Andreessen Horowitz named the trend LLMflation.

The per-token price is falling. The bill is not. Reasoning tokens, agentic workflows, context inflation and a subsidised inference floor all push the other way, which is why enterprise AI spend rose through 2025 even as unit prices collapsed.

The structural point is where the work happens. On our benchmarking, a decision-grade question answered against the structured record costs on the order of 5,000 tokens. The same question put to an unaided model re-reading raw scans runs from about 100,000 tokens into the millions. The multiple is illustrative rather than an audited constant, and it should be read that way, but the direction is not in doubt: normalisation is paid once, and re-reading is paid every time anyone asks anything.

Where should AI stop and the geologist start?

There is a new AI geology startup every week promising to tell you where to drill. Our view is that geology is interpretive in a way that does not yield to that yet, and that the companies furthest along this road have mostly concluded they would rather stake their own ground than sell the prediction.

So Pulse is infrastructure, not a target generator. We are not trying to hijack a geologist's thought process. The job is to compress the search to something close to zero, and leave the interpretation where it belongs. Give the geologist the questions, not the answers.

Chris put the counter-case on the call, and it is the right worry to hold onto. A model has never experienced the natural world, only the digital record of it. A geologist can pick up a sandstone, rub a thumb across it, and learn something no dataset carries. The risk is not that AI gets the rock wrong. It is that a geologist stops thinking deeply because the sweep looked complete.

Eight questions to ask before you trust a geological dataset

Run these on any record, any archive, and on anything an AI hands you from one.

#QuestionWhy it matters
1What was the original unit and basis, and who converted it?A converted number and a transformed number are different things.
2What datum and projection is this on?Relabelled is not reprojected. It is the most common silent error.
3Is the location a coordinate, or a matched place name?Matched locations need a confidence, not a pin.
4Can you see the source page, in the original language?If you cannot reach the page, you cannot defend the number.
5What did the process reject, and can you see it?What was discarded matters as much as what was kept.
6Is this the current statement or a superseded one?A right number from the wrong vintage is a wrong number.
7Does it compare across the set, or only within one report?Internal consistency is not comparability.
8Could you defend it in a technical review?If you would re-check it by hand, it has not saved you the work.

Get the technical white paper

Stone Age to Space Age: the hidden discipline of mining-grade data normalisation runs to 35 pages and twelve sections, written for practitioners rather than for finance. It carries the six normalisation problems in full, the scored pilot above with its methodology, the georeferencing proof, the cost curve, and the self-diagnostic.

Or read what the white paper covers first.


About Pulse Intelligence

Pulse Intelligence is the technology and infrastructure layer for mining data. We digitise, validate and structure NI 43-101, JORC and SK-1300 technical reports, exchange disclosure, government archives and historical geological records, then serve them through a platform, a documented API and the Pulse MCP, so the work runs inside your own model, your own GIS or your own systems rather than in a window you copy from. More than 450 investment and audit grade metrics across 29,600+ mining assets, every figure opening the original page of the original document in one click.

Less searching. More strategising.™

AI Readiness Diagnostic

Where does your team's data infrastructure sit today?

Answer 10 questions. Get a private diagnostic on your AI readiness — in minutes.

Pulse Intelligence

Less Searching. More Strategising.™

See the platform running on real mining data. Book a demo to see what this looks like for your team.