MIRAmodular interoperable research attribution Try it

Context and caveats

How MIRA was made

The usual workshop builds a schema. This one arrived with a converged schema — a year of user-story and interoperability calls behind it, published as a PR-able repo before the in-person phase — and spent its days implementing it across real tools and trying to break it on real data.

MIRA didn’t begin at the workshop. It generalizes a discourse-graph grammar proven over a five-plus-year deployment of Discourse Graphs in working labs, and it was funded on one condition — deliver a pilot that moves a real result from one tool to another, not another standards document. Through spring 2026, practicing scientists, tool-builders, funders and publishers turned their own workflows into the requirements; by April–May the group had converged on the six node types on this site — Question, Claim, Evidence, Study, Protocol, Request — with typed, attributable edges, and settled on JSON-LD over KOI as every tool’s shared endpoint.

Crucially, the schema went public as a PR-able LinkML repo before the in-person workshop began — held in a collectively owned org so every tool, Discourse Graphs included, inherits one common schema rather than any single tool’s. At the mira.science workshop in Ireland the test was never “is this a tidy ontology,” but whether the tools in the room could read and write it, and move a record between two of them. Both happened — and mid-demo, a subgraph published from one tool appeared in another. What shipped, with honest status labels, is on Try it.

Converged, not decreed; tested, not asserted; gaps tracked, not hidden. What the workshop didn’t close is filed the way the schema says work should be filed — in the open, as Request nodes and schema issues: governance, per-node attributes, endorsement, licensing.

Human-in-the-loop, placed precisely

Extraction is AI-assisted; commitment is human.

The hand-correction points aren’t an apology — they’re the interface. The places where a human decides what to keep and how strong a claim is are the point, and they should be explicit and visible in the data.

Every extracted node ships at a curation status and climbs a ladder: Initial AI draft → In expert review → Expert-verified. AI always starts a node at the bottom; only a human advances it. The status is surfaced as a per-node badge and a topology filter, and every AI-drafted node is anchored to a verbatim source quote, so the reviewer can check it in seconds rather than re-reading the paper. This maps onto machinery we already have: the Discourse Graphs NodeFormality ladder and MIRA-extraction’s provenance.excerpt.

Isn’t a human step a bottleneck? Only if it’s in the wrong place. Extraction, anchoring and structural validation scale with the models and get better every generation. What does not scale away is commitment — deciding what enters the shared record and how strongly it is claimed. We put the human exactly there, and nowhere else, and we make the boundary visible in the data rather than leaving it implicit. That turns a supposed weakness into the answer to the only question that matters for a shared record: who is accountable for this?

The pattern is running in public in the oasisresearchlab/language-and-health-open-synthesis graph — its three-tier curation status is live and filterable. See it on Try it.

How this entry was made

Written by a human and an AI — with the credit traced.

This submission was authored by Matt Akamatsu with Claude (Anthropic) over a short, intensive sprint. In the spirit of the schema — where every record points at its provenance — we logged that collaboration and distilled it into a public attribution table.

The short version: the MIRA schema and its community-tested grammar predate this work and underlie all of it — the entry argues for that structure, it didn’t invent it. The human supplied the steering: the thesis corrections, the case choices, the format pivots, the visual direction, and the guardrails — often as a single redirecting sentence. The AI did the investigation, turned those steers into drafts and build-ready specs, and wrote and debugged essentially all of the code. The two contested cases’ structural findings were AI investigation over public primary sources, chosen and framed by the human.

The full breakdown — idea by idea, with quotes — is the provenance & attribution table on GitHub.