The EQUORA Trilith Method™ claims that adversarial verification between different AI models catches errors that a single model misses, and that logging every machine contribution makes those errors attributable afterwards. Both claims are testable. This programme treats the method as an object of study rather than a settled framework, and it uses the institute's own published output as the test corpus.
Active · Method v1.0 documented · Protocol in developmentAI-native research organisations are appearing faster than the methods to evaluate them. The failure mode the Trilith Method™ is designed against is specific: an organisation that generates research at machine speed can generate convincing error at machine speed, and language models produce output that reads as persuasive whether or not it is true. A second, quieter failure mode is epistemic monoculture — if every researcher leans on the same few foundation models, the organisation inherits their blind spots in lockstep.
The method's answer to both is architectural: distrust is built in, model diversity is a design requirement, verification is adversarial by construction, and every machine contribution is logged so it can be audited later. What the method does not yet have is evidence that this architecture performs better than the alternatives it rejects.
The research gap: Framework proposals for verifiable AI-assisted research have begun to appear, and one of them states this programme's first hypothesis as a formal claim. Traxia (Dogah, arXiv:2606.08256) specifies an agent-native publishing infrastructure whose Design Property 6.2 asserts that layered adversarial review detects at least one genuine flaw with strictly greater probability than single-tier review at matched review effort — the same proposition as H1 below, argued from a partition of the flaw space rather than measured. That paper states plainly that it reports no empirical results and defers validation of the assumption to future work. The gap is therefore evidential rather than conceptual: the architectures now exist, and the evidence that they outperform the alternatives they reject does not. Existing work on LLM self-consistency, ensembling and self-critique measures agreement between outputs, which is a different quantity from error detection.
A measurement programme run by the party whose method is under test has attack surfaces of its own. Naming them is part of the design rather than an appendix to it, and two of them have no resolution at present.
| Surface | The case | Response |
|---|---|---|
| Governance capture | The same person sets the evidentiary bar and decides what is published. Layer III concentrates both, which is what makes it non-delegable and also what makes it a single point of failure. | The Research Governance Network is the intended structural answer. Until it operates, the neutrality of the governance layer is a stated assumption rather than a demonstrated property, and should be read as one. Open. |
| Selective correction | Corrections are recorded by the same party that made the error. Nothing forces disclosure of an error that no outside reader noticed, so the published correction log understates the true rate by an unknown margin. | The blind attribution study brings external readers into the record, and the corrections channel is open to anyone. Neither makes the log complete. Open. |
| Trace fabrication | A disclosed reasoning trace may be a reconstruction rather than a record — the trace fidelity problem. | H5 measures it directly instead of assuming it away, using an independently recorded route as the comparison. |
| Corpus selection | The seeded-error corpus is assembled by the party whose method it evaluates, and error placement can favour the protocol being defended. | The corpus, the seeding procedure and the scoring are published so that other groups can re-seed and re-run on their own model stacks. A result that only survives our seeding is not a result. |
| Metric gaming | Optimising for trace completeness produces exhaustive but vacuous steps that satisfy the form without carrying inference. | Steps are scored for informativeness as well as presence, and the corpus includes items whose correct handling requires a short trace. |
| Correlated models | Model diversity is assumed to decorrelate error. Shared training corpora and convergent post-training conventions may defeat that assumption. | H2 tests it. If diversity fails, the finding is reported and the weight shifts to the governance layer, with the Method's stated reason revised accordingly. |
This is not a claim that the Trilith Method™ has been validated. It is the plan by which it could fail. A method that separates functions by trust level, and then never checks whether the separation earns its cost, would be a stance rather than a method.
The programme also does not measure the governance layer's judgment, because judgment is the part deliberately kept out of measurement. What it measures there is documentation: whether a reader can find where a person decided, and on what basis.
This is a research plan rather than a result. It sets out what the programme intends to test, on what basis, and what would count as failure. Parts of it will turn out to be wrong; where that happens, the correction is recorded here with its date instead of being quietly removed.
Work by third parties is attributed to its sources and described at the confidence its evidence supports. Nothing on this page should be read as professional advice in the programme's domain, and nothing here has been peer reviewed unless a specific publication is cited as such.
Pölö (László Papp) — Founder, EQUORA Institute. Researchers working on LLM evaluation, research integrity or multi-model verification who would like to run the benchmark on their own model stack: lpapp@equora.institute