I was disappointed to wake this morning to find that my access to the Mythos Class of AI models had been suspended by an export-control directive issued by the United States Government.

That disappointment was not merely about losing access to a powerful model. It was about losing access to a capability class I had only just begun to test.

Most of the attention around frontier agentic AI has focused on software: autonomous coding, automated delivery pipelines, large-scale refactoring, enterprise systems migration, and the kind of infrastructure work that ordinarily requires coordinated engineering teams.

The examples are striking. A model takes a single product brief and returns a working full-stack application. It migrates a legacy codebase. It generates tests, debugs failures, repairs CI pipelines, writes documentation, and iterates until the build passes. It does not merely answer a question; it plans, delegates, executes, checks, and revises.

That is a profound shift.

But I wanted to test a harder question: whether the same capability could be deployed outside software — in a specialised reasoning discipline where the output is not code but judgement, exercised autonomously, and without having to be coddled by a human at every turn.

The Problem

There is a well-founded fear of AI in legal work. Used carelessly, AI is not merely unreliable; it fails fluently and convincingly. A nervous junior solicitor at least has the good grace to hesitate before citing a case that does not exist; the model supplies it with serene confidence, correct citation format and all.

A hallucinated case looks like a real case. A misstated principle looks like doctrine. The surface form is often immaculate, even when the substance is wrong.

Consider researcher Damien Charlotin’s database, which now records hundreds of legal decisions involving AI-hallucinated content — recent reporting places the figure above 700 court decisions, with roughly nine in ten handed down in 2025 alone.

That is not a marginal defect to be patched in the next release. It is a category of professional risk in its own right — one that attaches to the practitioner, not to the tool.

So the question is not whether AI can produce something that sounds legal. It plainly can. A little Latin, some well-placed legalese, a citation rendered in immaculate format — and the artefact looks the part. The harder question is whether that surface can be made to track the truth.

The Old Framework

Over the past few months, I had been building a customised legal-AI toolkit using earlier frontier models — primarily Opus 4.8 and GPT 5.5.

The goal was not generic legal drafting. It was a controlled legal-reasoning and drafting environment.

I connected the models to corpora of real legal documents, original judgment texts, and authoritative legal databases. I built citation-verification workflows; structured drafting conventions; register controls intended to hold a consistent voice; and simulated adversarial pressure-testing, in which the system was made to argue against its own output.

Legal reasoning proved more difficult. Simply prompting the use of the IRAC method taught in many Australian law schools was insufficient. Over many sessions, iterations, and manual reconfigurations, I developed a system that applied neurosymbolic logic to legal reasoning. It sought to combine the natural-language processing of large language models — the neural layer — with the rigid, rule-based logic of formal legal codes — the symbolic layer — to reduce both hallucination and black-box reasoning.

It could produce serious legal argument. It could reason through procedural issues. It could verify authorities when directed properly. It could approximate my drafting register with a fidelity that would have seemed implausible a year ago.

But it remained brittle.

Register drift persisted. Evidence became contaminated. Cross-references broke. Legal propositions were, at times, overstated. Most seriously, material issues were missed at an alarming rate. The system “worked” in a basic sense, but it worked because I was constantly correcting it — refining prompts, adjusting context, rewriting skills, rerouting workflows, and catching failures after they had appeared.

Then Came the Mythos Class

The Mythos generation represented something different.

I had heard the software use cases: multi-day autonomous tasks, ready-to-ship systems from a single prompt, 3D modelling of humanoid robotics, and the rest. I wanted to know whether the same pattern could be brought to legal reasoning.

The Experiment

I began with a single, detailed prompt that prescribed the fixed features of the approach: research, simulated testing under controlled conditions, looped iteration, and clear success criteria. Testing was to be rigorous and unforgiving. I instructed the use of mutation tests: fabricated citations, real authorities cited for the wrong proposition, truncated quotations, arithmetic traps, procedural-power traps, and unsupported factual assertions. The brief was specific and exacting; within the non-negotiables, however, I authorised the model to exercise its own discretion.

From that single prompt, the Fable model carried out a multi-day, autonomous task without requiring one further instruction from me. Its approach ran as follows.

Problem identification, analysis, and research. The research ran in two parts. The first was a general identification of the failure modes of AI in legal work — hallucinated authority, register drift, contaminated evidence, and missed issues — together with the methods proposed to contain them. The second was a granular critical analysis and audit of my existing framework: a mechanism-by-mechanism examination of where the neurosymbolic toolkit failed, and why, so that each weakness was traced to a specific cause rather than treated as one diffuse fault.

A new framework established. From that analysis it specified an architecture rather than a prompt: a set of components sitting over a shared, inspectable state.

Experimental parameters and conditions designed and enforced. It fixed the success criteria in advance and built the conditions under which the framework would be tested. Mechanical screens — deterministic checks — were erected as hard gates: any output that failed a screen did not proceed. Evaluation was blinded wherever judgement, rather than a deterministic rule, was the measure of success.

Simulated testing. It ran the framework against the full mutation battery, and against an instantiated opposing case constructed to attack its own arguments.

Assessment of results and identification of problems. It logged each defect, traced it to the component responsible, and isolated the cause rather than the symptom.

Looped testing and reiteration until the success criteria were met. It repaired, re-ran, and repeated — promoting each improvement one mechanism at a time.

The finished framework had eight components: orchestration, rule formalisation, research, argument graphing, drafting, verification, adversarial simulation, and benchmarking. Beneath them sat a state spine — a case file, a rule register, an authority ledger, an evidence map, an argument graph, and a paragraph trace.

What I was hardwiring, at last, was a set of standing rules. Every authority must be grounded before it is used. Every rule must be decomposed before it is applied. Every material factual assertion must trace to evidence. Every argument must be tested against an instantiated opposing case. Every draft must pass deterministic screens wherever deterministic checks are possible. Every evaluation must be blinded wherever judgement is involved. Every defect must be logged. And every improvement must be promoted mechanism by mechanism, on evidence — never on the strength of a good feeling.

Access to the model that powered the new system is now gone.

But the work it produced remains: a method, a benchmark, a defect log, and a clearer view of what agentic AI may become in specialised reasoning work.

What Transfers

A method, unlike a model, cannot be withdrawn by directive — and this one was never really about law.

What the framework encodes is a general discipline for any domain in which the work product is judgement, the cost of fluent error is high, and the reasoning must survive scrutiny after the fact. Strip out the legal vocabulary and what remains is a procedure: ground every claim before it is used; decompose every rule before it is applied; trace every material assertion to evidence; test every position against its strongest opponent; screen mechanically where a deterministic check exists; blind the evaluation where it does not; log every defect; and improve one mechanism at a time. None of that is peculiar to litigation.

Three fields make the point.

Clinical informatics. A decision-support tool that fabricates a contraindication, misreads a guideline, or quietly drops a comorbidity fails in precisely the way a hallucinated authority fails — plausibly, and at the point of greatest consequence. The same architecture answers it: the authorities, here the guidelines and the evidence base, grounded before use; assertions traced to the record; deterministic screens on whatever is checkable, such as dosage, interaction, and units; blinded review where clinical judgement is irreducible; and a defect log that converts each near miss into a permanent correction rather than a remembered anecdote.

Quality assurance and accreditation. Quality systems live or die on the auditable trail — the capacity to show not merely that a conclusion was reached, but how, on what evidence, and against which standard. A reasoning system built to trace every assertion to its source and to log every defect is, in effect, an accreditation instrument: it produces the evidence map and the corrective-action record as a by-product of doing the work, rather than reconstructing them under audit pressure after the fact.

Business transformation. Change programmes fail less often for want of ideas than for want of disciplined reasoning under uncertainty — assumptions left untested, options never put against their strongest counter-case, decisions taken on the strength of a good feeling and unrecoverable six months later because no one recorded why. The same spine applies: decompose the problem, ground each assumption, instantiate the opposing case, screen what can be screened, and keep a defect log of what went wrong and what was changed in response. The output is not a slide deck. It is a defensible decision trail.

The common thread is accountability. In each field the danger is identical — a confident answer that cannot be checked — and the remedy is identical: a system that cannot reach a conclusion without showing its working.

The future is not AI as an oracle — legal, clinical, or managerial.

It is AI as a disciplined reasoning system, held to account at every step.

And that is far more interesting.

The workflow, at a glance

One instruction → a six-stage research workflow

A single brief set the model running for hours: it did its own research, built a tool, then tested that tool like a scientist — against a control, with the judging blinded. Select a stage to expand it.

1
Parallel subagentsResearch

Four assistants worked in parallel: studied 67 real court submissions; audited the existing toolkit; read the academic literature on making AI reason like a lawyer; and profiled the author’s own writing voice from filed documents.

2
Parallel subagentsShared spec

Seven more assistants each wrote one part of a new eight-skill legal "suite" — its reasoning, research, drafting, fact-checking and self-criticism — to a shared design, plus a small program that mechanically checks every draft.

3
Verified fact packsNo contamination

Built four realistic but fictional legal cases, each with a sealed pack of pre-verified facts and authorities — so that any invented or misquoted source could be caught automatically. These are the controlled test conditions.

4
Identical briefsBlinded arms

Each case was drafted twice from identical instructions — once by the new suite, once by the existing toolkit (the "placebo control"). Neither side was told it was in a contest, so neither could try harder for the test.

5
AnonymisedOrder-swappedEvidence-quoted

Independent AI judges compared the two drafts with the author hidden and the order swapped — so a verdict only counted if it held up both ways. Every judgement had to quote the wording that justified it.

6
Correction loopNo result overstated

The judges' findings became a fault list. Nine faults were fixed and the weakest cases re-run and re-judged. The final verdict was reported straight: the new tool won on some tasks, the old one held its ground on others.

Fix & re-test — faults found at stage 6 were fed back into stage 4, and the weakest cases re-run.

What made this new

Not the answer — the method. The model planned a multi-day project, delegated to roughly twenty assistants, ran a controlled, blinded experiment on its own work, corrected itself, and refused to overclaim. That is research conduct, not a chat reply.

~20 AI assistants · 2 test rounds · 8 drafts judged · 9 faults fixed · live toolkit left untouched