Research Plan · Research Infrastructure · 2026–2028

Trilith Method™:
testing a method on itself

The EQUORA Trilith Method™ claims that adversarial verification between different AI models catches errors that a single model misses, and that logging every machine contribution makes those errors attributable afterwards. Both claims are testable. This programme treats the method as an object of study rather than a settled framework, and it uses the institute's own published output as the test corpus.

Active · Method v1.0 documented · Protocol in development
The Problem

A method that has never been measured

3
layers separating machine functions by how far the machine is trusted
11
research functions covered across the generative, verification and governance layers
3
independent model families required for triangulation, because architectures fail differently

AI-native research organisations are appearing faster than the methods to evaluate them. The failure mode the Trilith Method™ is designed against is specific: an organisation that generates research at machine speed can generate convincing error at machine speed, and language models produce output that reads as persuasive whether or not it is true. A second, quieter failure mode is epistemic monoculture — if every researcher leans on the same few foundation models, the organisation inherits their blind spots in lockstep.

The method's answer to both is architectural: distrust is built in, model diversity is a design requirement, verification is adversarial by construction, and every machine contribution is logged so it can be audited later. What the method does not yet have is evidence that this architecture performs better than the alternatives it rejects.

The research gap: Framework proposals for verifiable AI-assisted research have begun to appear, and one of them states this programme's first hypothesis as a formal claim. Traxia (Dogah, arXiv:2606.08256) specifies an agent-native publishing infrastructure whose Design Property 6.2 asserts that layered adversarial review detects at least one genuine flaw with strictly greater probability than single-tier review at matched review effort — the same proposition as H1 below, argued from a partition of the flaw space rather than measured. That paper states plainly that it reports no empirical results and defers validation of the assumption to future work. The gap is therefore evidential rather than conceptual: the architectures now exist, and the evidence that they outperform the alternatives they reject does not. Existing work on LLM self-consistency, ensembling and self-critique measures agreement between outputs, which is a different quantity from error detection.

The Object of Study

Three layers, and what each one has to prove

I.
Generative
Literature discovery, experiment design, simulation, knowledge graph, first drafts. The claim here is a speed claim, and it is the easiest of the three to demonstrate. It is also the layer that produces the errors the other two have to catch.
II.
Verification
The assert–refute–adjudicate protocol, multi-model triangulation, the research log, reproducibility checks. This is the layer the whole method rests on, and the one with no published evidence behind it. Most of this programme addresses it.
III.
Governance
Problem selection, the evidentiary bar, what goes out under the institute's name. Deliberately non-delegable, and therefore not a performance claim at all. What can be measured here is whether the human decision points are documented well enough that an outside reader can locate them.
Hypotheses

What we are testing

H1 — Adversarial verification outperforms single-model review
On a corpus of research claims with deliberately seeded errors, a structured assert–refute–adjudicate protocol across three independent model families detects a measurably higher proportion of seeded errors than review by any one of those models alone, at matched compute. Primary outcome: detection rate by error class. The protocol fails this test if the gain disappears once the single-model baseline is given the same number of passes.
H2 — Model diversity reduces correlated error
Errors that survive verification are less correlated across different model architectures than across repeated runs of the same architecture. If this holds, model diversity is doing the work the method attributes to it. If errors turn out to be correlated across architectures — because the models share training data and post-training conventions — then diversity is a weaker safeguard than the method assumes, and the governance layer carries more weight than currently stated.
H3 — Provenance logging makes surviving errors attributable
For a published claim later found to be inaccurate, the research log permits an independent reader to identify which layer introduced the error and which step failed to catch it. Outcome measure: the proportion of documented corrections for which an outside reader, given the log alone, reaches the same attribution as the internal review.
H4 — Error classes are stable enough to design against
The errors that survive verification fall into a small number of recurring classes rather than being idiosyncratic. Early live operation suggests candidate classes: compression of a correctly cited secondary source into an incorrect claim; quantitative assertions with no retrievable source; explanatory mechanisms carried over from popular literature without an evidentiary base; and orthographic or terminological drift in names and dates. If the taxonomy is stable, each class can be given its own check, which is the basis of the forthcoming Trilith Protocol.
H5 — A disclosed reasoning trace may not be the trace that was followed
H3 asks whether a log lets an outside reader locate an error. It presupposes that the log records what actually happened. Traxia (Dogah, arXiv:2606.08256) names the residual as the trace fidelity problem: no external verifier can guarantee that a disclosed trace is the one executed rather than a plausible reconstruction assembled to satisfy the disclosure requirement. That paper states it does not solve the problem, only raises the cost of fabrication, and that a full solution would require cryptographic commitments made at inference time by infrastructure that current deployments do not expose. This programme treats fidelity as measurable rather than assumed. The test uses derivations whose route is recorded independently, at the time of the work, by a channel separate from the one that later produces the disclosed trace; the two are then compared blind and divergence scored by class. Outcome: the proportion of traces that misdescribe the derivation, and whether the misdescription is detectable from the trace alone. If fidelity turns out to be low and undetectable, the Method's logging claim narrows to something worth stating plainly — that the log evidences the output rather than the reasoning — and H3 inherits that limit.
Methodology

Research and measurement approach

Seeded-error benchmark
A corpus of research claims drawn from the institute's own domains, into which errors of known class and severity are inserted. Each item is passed through the assert–refute–adjudicate protocol and through single-model review at matched compute. The corpus, the seeding procedure and the results are published so that other groups can run the same benchmark on their own model stacks.
Live error taxonomy
Every correction to a published EQUORA Institute or iterators.org page is recorded with its date, the claim as it stood, the corrected claim, and the layer at which the error entered. This produces an error taxonomy from real operation rather than from simulation, and it is the empirical base for H4. The corrections stay visible on the pages themselves.
Correlation analysis across architectures
The same claims are evaluated by model families with different training lineages, and the residual errors are compared for overlap. The analysis distinguishes errors that every architecture misses from errors that only one misses, which is the quantity that determines how much protection diversity actually provides.
Blind attribution study
For H3, readers outside the institute are given the research log for a documented correction, without the internal review, and asked to identify where the error entered. Agreement between external and internal attribution measures whether the log is genuinely auditable or only nominally complete.
Negative results as first-class records
Refuted hypotheses and failed approaches are stored in the knowledge graph alongside confirmed findings, attributed and searchable. This is both a methodological commitment and a research object: whether a negative-results store measurably reduces repeated dead ends is itself testable over the programme's duration.
Protocol specification
The Method is the stance; the Trilith Protocol is the executable specification — how an assertion, a refutation and an adjudication are run across models and recorded. The Protocol is drafted from the benchmark and taxonomy results rather than in advance of them, so that each prescribed check answers a documented failure.
Trace fidelity study
For H5, a subset of work is instrumented so that the route to a claim is recorded independently, at the time of derivation, by a channel separate from the one that later produces the disclosed trace. The two are compared blind. This is the only component of the programme that requires instrumentation before the fact rather than analysis after it, and it therefore constrains the schedule: fidelity cannot be established retrospectively.
Correction propagation
A dated correction currently stops on the page where the error was found. Where a corrected claim has been relied on elsewhere, the dependent pages are flagged and reviewed, so that a page resting on a withdrawn claim carries a visible marker rather than continuing to read as sound. The mechanism is adapted from the staleness score in Traxia, where the impact of a retraction propagates along the provenance chain.
Gaming resistance
Any measure of trace completeness invites Goodhart's law: a system optimising for the measure can produce exhaustive but vacuous steps that satisfy the form without carrying inference. The benchmark therefore scores steps for informativeness as well as presence, and the seeded corpus includes items whose correct handling requires a short trace, so that length alone cannot score well.
Read the Trilith Method™ framework →
Timeline

Research roadmap

2026 Q2
Method documented. Trilith Method™ v1.0 written up: three-layer architecture, eleven functions, boundary definition. Published with a persistent identifier for priority and attribution — all versions resolve through 10.5281/zenodo.21700279, currently v1.2.
2026 Q3–Q4
Live taxonomy begins. Correction logging introduced across iterators.org and the institute's research pages, with propagation to dependent pages. First error classes recorded from operation. Benchmark corpus assembled and seeding procedure specified. Independent route recording instrumented for the H5 fidelity study, which cannot be done retrospectively.
2027 Q1–Q2
Benchmark run. H1 and H2 tested on the seeded corpus across three model families. Results published whether or not they favour the method. Blind attribution study for H3 conducted with external readers, and the H5 trace fidelity comparison run against the recorded routes.
2027 Q3–2028
Protocol release. Trilith Protocol drafted from the results and released openly for replication. Preprint on Zenodo; submission to a venue in research methodology or AI evaluation. The benchmark is left runnable by other groups on their own stacks.
Threat Model

Where this programme could be fooled, including by us

A measurement programme run by the party whose method is under test has attack surfaces of its own. Naming them is part of the design rather than an appendix to it, and two of them have no resolution at present.

SurfaceThe caseResponse
Governance capture The same person sets the evidentiary bar and decides what is published. Layer III concentrates both, which is what makes it non-delegable and also what makes it a single point of failure. The Research Governance Network is the intended structural answer. Until it operates, the neutrality of the governance layer is a stated assumption rather than a demonstrated property, and should be read as one. Open.
Selective correction Corrections are recorded by the same party that made the error. Nothing forces disclosure of an error that no outside reader noticed, so the published correction log understates the true rate by an unknown margin. The blind attribution study brings external readers into the record, and the corrections channel is open to anyone. Neither makes the log complete. Open.
Trace fabrication A disclosed reasoning trace may be a reconstruction rather than a record — the trace fidelity problem. H5 measures it directly instead of assuming it away, using an independently recorded route as the comparison.
Corpus selection The seeded-error corpus is assembled by the party whose method it evaluates, and error placement can favour the protocol being defended. The corpus, the seeding procedure and the scoring are published so that other groups can re-seed and re-run on their own model stacks. A result that only survives our seeding is not a result.
Metric gaming Optimising for trace completeness produces exhaustive but vacuous steps that satisfy the form without carrying inference. Steps are scored for informativeness as well as presence, and the corpus includes items whose correct handling requires a short trace.
Correlated models Model diversity is assumed to decorrelate error. Shared training corpora and convergent post-training conventions may defeat that assumption. H2 tests it. If diversity fails, the finding is reported and the weight shifts to the governance layer, with the Method's stated reason revised accordingly.
What this programme does not claim

Boundary of the work

This is not a claim that the Trilith Method™ has been validated. It is the plan by which it could fail. A method that separates functions by trust level, and then never checks whether the separation earns its cost, would be a stance rather than a method.

The programme also does not measure the governance layer's judgment, because judgment is the part deliberately kept out of measurement. What it measures there is documentation: whether a reader can find where a person decided, and on what basis.

Method & Scope

How this page was produced

Trilith Method™ · Research Card
Programme status
Active · Method v1.0 documented · Protocol in development
Machine contributionI
Literature discovery and synthesis, candidate identification, computational modelling, and first drafts of this page. Volume and speed are the machine's contribution; none of it is treated as verified on its own.
VerificationII
Claims traced to primary sources rather than to summaries of them. Contested claims run through assert–refute–adjudicate across independent model families. Errors found after publication are corrected on this page with their date.
Human governanceIII
Problem selection, the evidentiary bar, and the decision to publish rest with Pölö (László Papp), EQUORA Institute, who holds editorial responsibility for this page.
Evidence basis
Claim confidence: stated per component in the Methodology section, with preprints, pilot studies, single trials and modelled estimates named as such at the point of use. Replication: none of the hypotheses on this page has been independently replicated; this is a plan, and that is what a plan means. Trace completeness: the reasoning behind each claim is given in prose rather than as a machine-verifiable trace, and its fidelity is itself under test in H5.
Version
v1.0 · 2026-07-29
What this page is

This is a research plan rather than a result. It sets out what the programme intends to test, on what basis, and what would count as failure. Parts of it will turn out to be wrong; where that happens, the correction is recorded here with its date instead of being quietly removed.

Work by third parties is attributed to its sources and described at the confidence its evidence supports. Nothing on this page should be read as professional advice in the programme's domain, and nothing here has been peer reviewed unless a specific publication is cited as such.

Research Lead

Principal investigator

Pölö (László Papp) — Founder, EQUORA Institute. Researchers working on LLM evaluation, research integrity or multi-model verification who would like to run the benchmark on their own model stack: lpapp@equora.institute