Reconciliation as a high-value application of machine intelligence

Bright Meadow Group


Ask what artificial intelligence is best at and you get a capability list: drafting, coding, summarizing, image generation, customer contact. Useful, all of it, and mostly a description of what the technology does easily rather than where it is worth the most. The better question is where machine capacity meets a problem that is large, well-defined, unglamorous, and currently unserved. By that test the answer is not a product category. It is a maintenance job that nobody has been able to staff.

The scientific and technical record needs to be reconciled against itself. Machines can do this. Humans cannot, and have never been able to.


Observe

The written record of human knowledge has been accumulated under conditions that guarantee internal contradiction. Findings are published in domain-specific venues, read by domain-specific audiences, and cross-checked against neighboring domains only when an individual researcher happens to sit at the seam. Nobody is assigned to consistency. There is no office for it, no budget line, no career track.

The consequences are documented and unsurprising.

Reanalysis of published work from its own reported numbers turns up arithmetic and reporting errors at rates that would end a career in accounting. The many-analysts studies are worse: hand one dataset to thirty independent teams with one question, and effects come back pointing in opposite directions, every path defensible. Whole literatures sit adjacent to one another with implications neither has stated, because the two readerships do not overlap. Don Swanson demonstrated this in 1986 by connecting fish oil to Raynaud’s syndrome through two bodies of work that shared no citations and no readers. He called what he found undiscovered public knowledge. The category is real and the inventory is large.

None of this requires anyone to have behaved badly. It follows from the structure of the enterprise. Specialization is what made modern science productive, and consistency-checking across specialties is the cost that was never paid.

Machine reconciliation is the first mechanism capable of paying it.


The claim, stated plainly

A sufficiently capable system fed the record and enough compute will drive it toward coherence — cross-checking claims against every other claim it holds, surfacing contradictions, and forcing resolution.

This works against falsehood better than most people expect. A false claim is never a single statement. It is a statement plus every consequence it entails, and those consequences propagate into domains the author does not control. Sustaining one requires forging its entire causal neighborhood. The cost scales with the density of cross-checking, which is precisely the quantity a reconciler raises. Uncoordinated deception does not survive this. Neither does most motivated reasoning, which is deception with the intent removed.

The corollary matters more than the claim. Truth persists because it is the only configuration that fits everywhere at once. A lie can be repeated a hundred ways; it still has to slot into the puzzle, and eventually it does not.

The division of labor this implies is the design decision that makes the whole thing work. Humans generate ideas, anomalies, and new measurement. The machine keeps the books. It is not asked to discover. It is asked to enforce, at a scale and density no institution has ever managed, and to report what will not fit.


Design

Four properties separate a reconciler that works from one that produces confident nonsense.

The residual pile is the product. Any flattening operation forces pieces. Log every one. The residuals — claims that had to be pushed to make the surface smooth — are the discovery queue handed back to the humans. A system that reports only its reconciled output has thrown away the valuable half. Kepler’s eight arcminutes of Mars residual were legitimately absorbable; every colleague he had could have added one term and satisfied the math completely. He refused, and got the ellipse. Refusal of that kind is not in the data, which is exactly why the residuals have to leave the machine intact.

Residual shape is a diagnostic. Scattered randomly, residuals are noise. Clustered by instrument, laboratory, decade, or funding source, they are systematic error — and clustering is the only handle available on bias detected from the inside. Noise flattens well. Bias does not flatten at all, because bias arrives pre-coherent. That is what makes it bias rather than error.

Underdetermination gets held, not collapsed. Where evidence genuinely fails to distinguish between two accounts, the correct output is both accounts plus the measurement that would separate them. Ptolemaic astronomy is the standing warning: epicycles are a Fourier series, residuals can be driven arbitrarily low by adding terms, and fourteen hundred years of improving data made the fit better rather than triggering a crisis. Precision was never the problem. A reconciler that resolves ambiguity toward whichever side has more documents has reproduced the failure with better hardware.

Plurality over canon. Cross-validation works because sources are independent. One canonical reconciler has systematic errors invisible from inside itself — there is nothing left to disagree with it. Several, built independently against the same evidence, turn divergence into a measurement. Agreement is informative. Disagreement locates either an underdetermination or a bias, and both are worth finding. Open corpus, open method, open weights, fully auditable. A single reconciler is also a single thing to capture.


The failure this cannot fix from inside

Everything above operates on pieces that exist. The record’s deepest defect is the pieces that were never cut.

Null results go unpublished. Analyses that came out wrong get abandoned. Trials that disappointed their sponsor never reach a journal. A coherence engine reading that literature finds beautiful agreement on effects that are not there, because nothing contradicts them — the contradicting measurements were never made into documents. The puzzle metaphor assumes the pieces are a sample of the picture. They are a sample of what someone chose to cut.

There is exactly one repair, and it does not come from the corpus. It comes from proving that a commitment existed before the result did. Hash the protocol, the hypothesis, and the analysis plan; timestamp them on a ledger nobody can backdate. Silence becomes legible. An unfulfilled registration is a visible hole. Twenty registrations against three publications tells a reconciler the shape of the seventeen missing pieces.

This is the one function in the architecture that requires an untamperable ordered ledger and cannot be substituted with a repository under administration. Note what it is for: the commitments, not the data. The corpus itself wants content addressing and Merkle history — tamper-evidence, independent replication, and a legible chain of supersession where nothing is deleted and every replacement is auditable. That machinery is deployed, mature, and effectively free. Consensus mechanisms solve a different problem: who writes next when two writes compete for one slot. Data deposit has no such conflict. Two laboratories depositing contradictory results is the system working.


Intervene

Compute is not the binding constraint. Deposit is.

The technical layer of this program is largely solved and largely cheap. What gates it is whether people are made to hand over their numbers. Mathematics has already run the experiment: formal verification projects are machine-checking the corpus from the axioms upward, and it works, because in mathematics the data is the argument and the deposit problem does not exist.

For everything else, the interventions are unglamorous and institutional.

Instrument-level deposit as a condition of funding and of publication, from here forward. Retroactive deposit wherever raw data still exists — much of the twentieth-century record survives only as summary statistics, the underlying measurements lost to dead formats, closed filing cabinets, and retired investigators. That portion is unrecoverable, and pretending otherwise wastes effort better spent on the forward commitment. Standing funding for reanalysis, which currently has no career path and no money. Pre-registration with hashed timestamps as default practice rather than exception.

Every item on that list is a fight with publishers, with pharmaceutical sponsors, with institutional review boards, and with the incentive structure of academic promotion. It moves at the speed of policy rather than engineering, which is slower than anyone building the technical side would like.

That is the honest shape of the opportunity. The most valuable available application of machine intelligence is a maintenance function on a shared asset that has gone unmaintained since the asset was created. It requires no capability that does not exist. It requires that the measurements be made available, and that the residuals be published rather than smoothed.

Flattening is the mechanism. New measurement pushed into the record is the engine. A closed corpus, however well reconciled, only becomes more consistent with itself.

Spread the love

Related Posts