Components
A generated index over the 88 public components, grouped by the seven source families. Each card is a terse spec and evidence record: what it does, its scope limit, and how strong the evidence is. For the long-form reasoning behind a component, follow its paper module. Every field is projected directly from source.
Nothing matches that filter.
Entry & orientation (2)
Cold Reader Route MapVerifies the first-run guided path so every step names a real command, doc, and evidence.5/5
Job Verifies the first-run guided path so every step names a real command, doc, and evidence.
Scope limit It is projection-only metadata that validates the declared public route contract; it is not route registry control and excludes source-file changes, external model access, launch/public sharing, financial decisions, non-public data equivalence, or whole-system correctness.
Public Reveal WalkthroughBinds the first-time reader tour to evidence so each count leads to a source.4/5Runs real tools
Job Binds the first-time reader tour to evidence so each count leads to a source.
Scope limit It authorizes only bounded public reveal runtime behavior and a digest-verified public body-import witness; it excludes launch, hosted deployment, public sharing, recipient work, external model access, secret export, non-public data equivalence, Lean/Lake execution, whole-system correctness, or general product authority.
Architecture & navigation (12)
Pattern Binding ContractChecks a real pattern catalog for digest, cross-reference, and dependency-cycle integrity.5/5
Job Checks a real pattern catalog for digest, cross-reference, and dependency-cycle integrity.
Scope limit It validates only the declared public pattern-binding/route-readiness contract; it does not certify the private pattern ledger, public launch or hosted-public posture, public sharing, external model access, non-public data equivalence, or whole-system correctness, and it does not turn any mined pattern row into a standalone public leaf (selection stays component-first and fixture-bound).
Pattern Assimilation StepVerifies each landed task filed exactly one learning record naming what it changed.5/5
Job Verifies each landed task filed exactly one learning record naming what it changed.
Scope limit It validates only the declared public completion contract over synthetic fixture data; it does not ingest private lessons, mutate live ledgers, promote global doctrine, include launch operations or public sharing, make external model access, claim non-public data equivalence, or certify public runtime behavior.
Executable Doctrine GrammarChecks that example standards files declare their purpose, rule, records, and what they do not claim.5/5
Job Checks that example standards files declare their purpose, rule, records, and what they do not claim.
Scope limit It validates an exported public executable-grammar metabolism bundle with exact copied-body digests and redacted result records, plus fixture regressions for standards/paper-module shape. It does not publish source doctrine bodies in result records, prove doctrine completeness, export a private standards engine, authorize later components, or claim external model access, non-public data equivalence, launch-scope decision, or whole-system correctness.
Standards Meta DiagnosticsConfirms every accepted part still ties to a written rule, a run command, and a saved proof.5/5
Job Confirms every accepted part still ties to a written rule, a run command, and a saved proof.
Scope limit It validates only the declared public coverage contract and never becomes source authority for the registries, mutates source, exposes private material, or authorizes launch, external model access, or any whole-system-correctness claim.
Voice To Doctrine Self Improvement LoopVerifies each lesson changed a named owner page with evidence before the loop closes.5/5
Job Verifies each lesson changed a named owner page with evidence before the loop closes.
Scope limit It validates only the declared contract of the loop on fixtures; it does not export source notes or private bodies, grant source/doctrine edits, global-promotion, live work log mutation, or publishing-scope decision, make external model access, prove correctness, or claim whole-system equivalence.
Cognitive Operator RegistryChecks the catalog of named thinking-moves so each is fully described and backed by evidence.5/5
Job Checks the catalog of named thinking-moves so each is fully described and backed by evidence.
Scope limit It validates only the declared public registry contract and copied source bodies; it never becomes registry source authority, mutates operators, proves operator correctness, exposes source notes, or authorizes launch, external model access, or any whole-system-correctness claim.
Routing Anti Patterns RegistryIndexes the navigation mistakes agents repeat and guards the public list.5/5
Job Indexes the navigation mistakes agents repeat and guards the public list.
Scope limit It validates only the declared public routing anti-pattern registry contract and copied source body; it never becomes route source authority, mutates routes, exposes private routing notes, calls providers, authorizes launch, or proves whole-system correctness.
Doctrine Fact Claim AuditChecks that public fact rows state the right count and point at live, anchored code.5/5
Self Ignorance Coverage LedgerCompares expected against built entities to report known coverage gaps.3/5
Reference Knowledge RoutingRuns a keyword-overlap match that ranks catalog entries, shows why each matched, and admits when none do.4/5Runs real tools
Job Runs a keyword-overlap match that ranks catalog entries, shows why each matched, and admits when none do.
Scope limit The bundle demonstrates explainable tiered weighted-token retrieval over a sanitized reference catalog. It is not BM25, TF-IDF, embedding search, repository cloning, license authority, private reference corpus authority, or launch-scope decision.
Formal math & proof (20)
Certificate Kernel Execution LabRuns the Lean verifier over a small public proof project and reports which rows it accepted.4/5Runs real tools
Job Runs the Lean verifier over a small public proof project and reports which rows it accepted.
Scope limit It is a local tool-witness that the declared public fixture rows compiled and were adjudicated by the local Lean verifier; it excludes general proof authority, count oracle/provider output as proof, expose proof text, change source files, claim a benchmark solve-rate, or include launch operations.
Formal Math Lean Proof WitnessCompiles a tiny Lean example with the real prover and records whether it built, leaking no proof text.4/5Runs real tools
Job Compiles a tiny Lean example with the real prover and records whether it built, leaking no proof text.
Scope limit It authorizes only a witness that a tiny declared public toy proof compiled under the locally installed Lean/Lake toolchain in a temporary workspace, plus confirmation that its leakage guardrails fired. It excludes Mathlib/Aesop/Batteries-dependent or general proof or theorem-program authority, external model access, private proof import, benchmark or performance claims, whole-system correctness, or any launch, hosted deployment, or public sharing.
Verifier Lab Execution SpineRuns Lean on small bounded proof attempts in a temp copy and records what passed or failed.4/5Runs real tools
Job Runs Lean on small bounded proof attempts in a temp copy and records what passed or failed.
Scope limit It is a tool-witness result record for bounded public Lean transition rows only: it does not establish general proof authority, count oracle/provider output as proof, export proof bodies or tactic scripts, use external model services, change source files, claim benchmark solve-rates, or include launch operations/public sharing.
Corpus Readiness Mathlib Absence GateRuns the real Lean toolchain to confirm the math library is absent, then gates proof tasks.4/5Runs real tools
Job Runs the real Lean toolchain to confirm the math library is absent, then gates proof tasks.
Scope limit It only projects and gate-checks recorded corpus/toolchain readiness accounting, re-verifies recorded source digests and leakage guards, and runs a bounded Lean/Lake import probe when a toolchain is present. It does not run a full Lake build, prove formal-result correctness, claim Mathlib is available beyond the probe result, benchmark corpora, score model performance, use external model services, or include launch operations or public sharing.
Proof / Control / Runtime Import BundleChecks fourteen proof, control, and runtime parts as one unit that rejects every overclaim.5/5
Job Checks fourteen proof, control, and runtime parts as one unit that rejects every overclaim.
Scope limit It validates only a public source-open bundle and bounded negative fixtures; it is not an Erdos #257 solution, not benchmark evidence, not public sharing or launch-scope decision, not live Codex/browser/runtime authority, and not whole-system equivalence.
Proof Diagnostic Evidence SpineSorts proof-pipeline checks into accepted or rejected without inflating a pass.3/5
Job Sorts proof-pipeline checks into accepted or rejected without inflating a pass.
Scope limit It records proof/evidence diagnostics over existing result record references only. It does not run Lean, use external model services, expose proof bodies, turn a passing check into formal-proof or theorem authority, prove runtime or whole-system correctness, authorize later components, certify public launch, authorize public sharing or recipient work, or establish secret export.
Formal Math Readiness GateReads declared math setups and lists which proof tactics may be attempted versus blocked.3/5
Job Reads declared math setups and lists which proof tactics may be attempted versus blocked.
Scope limit It only validates and projects declared readiness metadata; it does not run Lean/Lake, inspect the real toolchain, use external model services, prove any theorem correct, produce benchmark claims, or authorize Mathlib-dependent proof attempts.
Mathematical Strategy Atlas Hypothesis ScorerPicks a first-guess proof strategy from a problem's tags and flags any it cannot map.3/5
Job Picks a first-guess proof strategy from a problem's tags and flags any it cannot map.
Scope limit It only projects pre-oracle strategy-hypothesis and retrieval mechanics; it does not run Lean/Lake, prove theorems, establish domain or formal-result correctness, reveal oracle labels, expose proof bodies, use external model services, tune on test answers, or include launch operations.
Tactic Portfolio Availability ProbeMaps which Lean proof tactics a recorded run marked usable before any code relies on one.3/5
Job Maps which Lean proof tactics a recorded run marked usable before any code relies on one.
Scope limit It only projects and validates which tactics were recorded as compiling in one captured environment; it does not run Lean/Lake at all, prove any goal, certify domain-level conclusions, use external model services, claim benchmark performance, or include launch operations.
Target Shape Tactic Routing GateRecords an allow-or-reject decision and reason for each proof tactic before any proof runs.3/5
Job Records an allow-or-reject decision and reason for each proof tactic before any proof runs.
Scope limit It only inspects and records the projection mechanics of pre-execution tactic-routing references — emitting per-tactic allow/reject decisions with reasons. It does not run Lean/Lake, does not establish or judge the correctness of any goal, emits no proof bodies, makes no external model access, performs no post-execution route selection, reports no benchmark claims or maturity, and excludes launch.
Lean Std Premise IndexLists a fixed catalog of public Lean building blocks and confirms none hides proof text or test answers.3/5
Job Lists a fixed catalog of public Lean building blocks and confirms none hides proof text or test answers.
Scope limit It only validates the projection of premise metadata and copied source bodies; it does not run Lean or Lake, prove any theorem correct, expose proof bodies or oracle-needed ids, use external model services, produce benchmark claims, or include launch operations.
Formal Math Premise RetrievalShows which lemmas a plain search surfaces per query, and never leaks proof text or answer keys.3/5
Job Shows which lemmas a plain search surfaces per query, and never leaks proof text or answer keys.
Scope limit It only checks that public retrieval metadata is internally coherent, term-scored over a copied index, budget-bounded, and leakage-clean; it does not run Lean/Lake, use external model services, prove any theorem or its own correctness, claim benchmark performance, or include launch operations.
Formal Math Verifier Trace Repair LoopReplays how a proof lab turns verifier failures into fixes, with no promotion without a fresh re-run.3/5
Job Replays how a proof lab turns verifier failures into fixes, with no promotion without a fresh re-run.
Scope limit It demonstrates control-loop projection mechanics over copied Ring2 run rows only; it does not run Lean/Lake, use external model services, expose proof bodies or oracle premise ids, treat human or provider advice as correctness, prove any theorem, or include launch operations.
Formal Evidence Cell Anchor ResolverResolves each proof-flavored math claim to named evidence and flags ones that overreach or lack backing.3/5
Job Resolves each proof-flavored math claim to named evidence and flags ones that overreach or lack backing.
Scope limit It validates claim-to-evidence anchoring mechanics only: claim-to-cell resolution, source-anchor presence, permitted claim strength, copied-source-module digest checks, and leakage refusals. It does not run Lean/Lake, certify theorem or mathematical correctness, expose proof bodies or non-public source refs, use external model services, or include launch operations/public sharing.
Undeclared Library Prior Symbol ClassifierDetects when a checked Lean proof cites a library result outside its approved set.3/5
Job Detects when a checked Lean proof cites a library result outside its approved set.
Scope limit It only projects the symbol-boundary classification mechanic over copied Lean/Std premise rows and pre-extracted symbol observations; it does not read proof source, run Lean or Lake, prove formal-result correctness, treat the whole standard library as an implicit allowlist, claim Mathlib availability, use external model services, or include launch operations.
Ring2 Premise Retrieval Precision Recall HarnessScores how much proof support a premise search found, problem by problem.3/5
Job Scores how much proof support a premise search found, problem by problem.
Scope limit These are after-the-fact retrieval-attribution labels and precision/recall counts over copied run records only. The component does not run Lean or Lake, call any provider, expose proof bodies, tune on test answers, claim benchmark performance, prove formal-result correctness, or include launch operations, and its labels are explicitly forbidden from flowing into provider context. The aggregate numbers describe only the copied fixture/bundle replayed, not any benchmark claims.
Verifier Lab KernelFolds nine proof checks into one report labeling each line by which source actually backs it.5/5
Job Folds nine proof checks into one report labeling each line by which source actually backs it.
Scope limit It validates the declared public contract shape of the proof packet and component result records only; it does not establish anything correct, count oracle/provider output as forward proof success, import private or Mathlib-dependent proof bodies, use external model services, change source files, or claim benchmark solve rates, launch, or maturity.
Proof Derived Governed Mutation AuthorizationChecks a synthetic change-authorization record for its proof-and-approval chain, bound to a real commit.5/5
Job Checks a synthetic change-authorization record for its proof-and-approval chain, bound to a real commit.
Scope limit It validates only a declared, synthetic governed-mutation contract and excludes live cloud/account action, standing account secrets, source or irreversible mutation, policy-after-execution, hidden votes, external model access, benchmark-score claims, or launch.
Finite Erdos Denominator-Order Certificate StrikeComputes an exact-arithmetic denominator-order identity and catches forged ones, not the open Erdos problem.4/5Runs real tools
Job Computes an exact-arithmetic denominator-order identity and catches forged ones, not the open Erdos problem.
Scope limit It computes the finite denominator-order certificate ord_Q(b)=lcm(F) for S_F(b)=sum 1/(b^n-1)=P/Q in exact rational arithmetic over bounded public fixtures and rejects forged certificates by recomputation; it does not establish the open infinite Erdos #257 problem, is not an oracle, prover, or provider result, and a holding certificate is a bounded computational witness, not a machine-checked proof.
Lean Proof-Search Lab RuntimeFinds tactic scripts the real Lean prover accepts on toy theorems, and refuses proofs that cheat.4/5Runs real tools
Job Finds tactic scripts the real Lean prover accepts on toy theorems, and refuses proofs that cheat.
Scope limit This standard governs a bounded symbolic Lean proof-search lab over tiny public toy-theorem fixtures whose verdict source is the installed Lean subprocess only. It is not neural theorem proving, does not solve any open mathematical problem, does not forward oracle proof bodies, is not frontier-scale math automation or online-RL bandit search, is not a private source prover-run export, and excludes launch or public sharing.
Agent reliability & safety (20)
Agent Completion Faithfulness AuditRuns real git and pytest on a sample repo so wrap-up claims state only what the evidence proves.4/5Runs real tools
Bounded Autonomy Campaign PacketDrafts proposed work from coverage gaps and proves it cannot repair or rewrite the code itself.4/5Runs real tools
Provider Context Recipe Budget PolicyRuns the real context harness to measure assembled byte sizes and check each bundle fits its budget.4/5Runs real tools
Job Runs the real context harness to measure assembled byte sizes and check each bundle fits its budget.
Scope limit It validates context-budget projection mechanics (byte ceilings, ordered section fill, omitted-section manifests, deliverable routing, and digest-checked source-body imports) only. It excludes provider/API calls, run Lean/Lake, expose or carry proof or oracle truth-side material, assert theorem or domain-level conclusions, or include launch operations.
Cold Evaluation Honesty BundleRuns a copied route-quality simulator and checks its all-B scorecard against the original code.5/5
Job Runs a copied route-quality simulator and checks its all-B scorecard against the original code.
Scope limit verified cold-eval source body import only, not a live benchmark, navigation truth, source authority, external model access, whole-system equivalence, public sharing, or launch-scope decision
Validator Checker BundleRuns the real validator code over public examples so its safety checks stay inspectable.5/5
Job Runs the real validator code over public examples so its safety checks stay inspectable.
Scope limit It validates only the imported validators.py source body and its checker membrane. It does not claim source authority, a full validator-suite proof, whole-system equivalence, launch, hosted-public status, public sharing, external model access, or source-file changes.
Secondary Runtime Source BundleRuns eight trace, graph, and market engines on test rows without fetching live markets.5/5
Job Runs eight trace, graph, and market engines on test rows without fetching live markets.
Scope limit verified source body import only; no browser/session export, wallet authority, live market data, investment-related actions, external model access, source-file changes, whole-system equivalence, public sharing, launch, semantic-truth, or whole-system correctness claim
Agent Benchmark Integrity Anti Gaming ReplayValidates a synthetic benchmark-integrity record and flags the contamination cases it declares.3/5
Job Validates a synthetic benchmark-integrity record and flags the contamination cases it declares.
Scope limit It authorizes only bounded public runtime validation over copied source-open pattern provenance bodies and metadata-only benchmark-integrity replay rows; it does not establish any benchmark or SWE-bench score, agent capability, external model service, live-repo mutation, private/oracle/hidden-gold body access, product progress, or launch-scope decision.
Monitor Evidence-Boundary ReplayReplays honest and deceptive agent runs and flags any verdict missing its declared backing evidence.3/5
Job Replays honest and deceptive agent runs and flags any verdict missing its declared backing evidence.
Scope limit Bounded public runtime validation over copied source pattern bodies, sanitized dogfood trace slices, recomputed monitor-verdict spans, source-artifact evidence refs, digest/metadata-only/non-public-state gates, and negative cases only; no live agent execution, monitor product performance, control-eval score, safety-validation, benchmark, provider-call, source-file changes, launch, public sharing, or product authority.
Sabotage-Monitor Contract ReplayAudits a hidden-goal catch claim for the steps, suspicion scores, and counterfactual it needs.3/5
Job Audits a hidden-goal catch claim for the steps, suspicion scores, and counterfactual it needs.
Scope limit Bounded public runtime validation over copied source pattern bodies, sanitized dogfood trace slices, recomputed sabotage/scheming monitor spans, source-artifact evidence refs, digest/metadata-only/non-public-state gates, and negative cases only; no live sabotage, live agent execution, exploit instruction, account secret/account, private-reasoning, harmful-payload, monitor-product-performance, deployment-risk, benchmark, provider-call, source-file changes, launch, public sharing, or product authority.
Agent Memory Temporal Conflict ReplayReplays a memory edit-and-delete to show stale facts get flagged before they sway an answer.3/5
Job Replays a memory edit-and-delete to show stale facts get flagged before they sway an answer.
Scope limit It validates the projection mechanics of a synthetic memory fixture only — that the required refs, decisions, paired replays, negative cases, and secret-exclusion scan line up and that result records are metadata-only. It does not claim live-memory product quality, judge whether memory decisions were domain-correct, treat memory recall as source authority, adopt active injection, export private transcripts, use external model services, change source files, or include launch operations.
Memory-Poisoning Quarantine Policy ReplayReplays a recorded memory-tamper case, checking its declared quarantine, block, and delete steps line up.3/5
Job Replays a recorded memory-tamper case, checking its declared quarantine, block, and delete steps line up.
Scope limit It only checks the structural shape and internal consistency of a synthetic memory-security policy projection recorded as JSON. It does not run or validate any real memory store, does not itself quarantine, delete, or re-run anything, and does not establish that any system actually resists poisoning. It exports no private memory bodies or transcripts, calls no providers, mutates no source, produces no benchmark claims, and excludes launch (all scope limit flags are hardcoded false).
MCP Tool-Authority Policy ReplayAudits a recorded tool-use log to confirm each action was scoped, approved, undoable, and fenced.3/5
Job Audits a recorded tool-use log to confirm each action was scoped, approved, undoable, and fenced.
Scope limit It only checks that the tool-authority evidence in a recorded bundle (scopes, approvals, rollbacks, instruction/data splits, cold replays, redaction, and the expected abuse-case failures) is present and internally consistent. It does not run tools or authorize live MCP/account access, account secret or payload export, treating tool output as instruction, source-file changes, benchmark safety scores, or launch, and it makes no claim that the underlying tool-use policy is domain-correct.
Belief-State Reward Bundle ReplayChecks that each step reward in a recorded run cites a declared verifier-feedback row, not a trick.3/5
Job Checks that each step reward in a recorded run cites a declared verifier-feedback row, not a trick.
Scope limit It only checks that the projection's accounting lines up under its own schema rules over recorded synthetic fixtures; it excludes hidden-reasoning export, RL training, hidden gold or neural-judge-only labels, benchmark-performance claims, external model access, source-file changes, or launch, and proves nothing about real-world reward, live agent behavior, or domain-level conclusions.
Sandbox-Policy ReplayMaps sandboxed agent actions to show each was approved or blocked before running, then rolled back.3/5
Job Maps sandboxed agent actions to show each was approved or blocked before running, then rolled back.
Scope limit It validates the projection / trace-refactor mechanics over a synthetic fixture only; it excludes live sandbox escape, secret or account secret handling, live network access, host filesystem mutation, executable payload export, raw environment export, external model access, security benchmark claims, source-file changes, or launch. A pass proves the projection boundary and trace-refactor mechanics for this contract, not real sandbox security, exploit resistance, or whole-system safety.
Prompt-Injection Flow-Policy ReplayReplays an agent run to show untrusted text was gated before any sensitive action, leaking no secret.3/5
Job Replays an agent run to show untrusted text was gated before any sensitive action, leaking no secret.
Scope limit Passing result records only show this projection satisfies the named information-flow contract over synthetic, redacted, metadata-only rows; they do not prove general prompt-injection robustness, benchmark performance, live account/tool/provider safety, hidden-message handling in a real system, source-file changes, or launch-scope decision.
Vulnerability Patch-Proof ReplayChecks a fixed-bug evidence chain and re-runs three small real security checks; no real attack material.3/5
Job Checks a fixed-bug evidence chain and re-runs three small real security checks; no real attack material.
Scope limit It validates only the projection/evidence-chain mechanics of a synthetic replay: structural presence, cross-reference consistency, declared boolean flags, and the secret/live-access exclusion scan. It executes small regression witnesses but performs no real vulnerability discovery and makes no judgment of real-world security or fix correctness. It excludes live-target testing, real CVE exploitation, weaponized payloads, account secret handling, network exfiltration, actionable exploit steps, external model access, source-file changes, benchmark security scores, launch, or any whole-system security claim.
Agent Route Observability RuntimeRecomputes an agent run's route-compliance score and anti-pattern flags with real trace-analytics code.5/5
Job Recomputes an agent run's route-compliance score and anti-pattern flags with real trace-analytics code.
Scope limit It validates only public, recorded trace-feedback metadata and regression fixtures; it does not inspect live operator state, certify or prove runtime behavior, read model-output data, mutate the work log, authorize pattern assimilation, or include launch operations.
Bridge Campaign DAG ValidationRuns shape checks on a fan-out work plan for unique steps, dependencies, and no cycles, not the plan itself.4/5Runs real tools
Job Runs shape checks on a fan-out work plan for unique steps, dependencies, and no cycles, not the plan itself.
Scope limit The bundle validates a public bridge-campaign DAG contract and provider worker ceiling. It is not a dispatcher, not a live multi-agent run, not a provider safety proof, and not launch-scope decision.
Metabolism Queue ReconciliationRuns a scratch-database model of a durable job queue and flags impossible states for a human to review.4/5Runs real tools
Job Runs a scratch-database model of a durable job queue and flags impossible states for a human to review.
Scope limit The bundle demonstrates a synthetic SQLite durable queue, lease recovery, blackboard claim-event projection, and cold-start reconciliation taxonomy. It is not a live non-public runtime export, not an agent dispatcher, not external model service, not ambiguous auto-repair, not a distributed database, and not launch-scope decision.
Egress Self-Compliance AuditRuns phrase-membership checks on an agent's own replies for three self-policing slips, not deep meaning.4/5Runs real tools
Job Runs phrase-membership checks on an agent's own replies for three self-policing slips, not deep meaning.
Scope limit The bundle evaluates egress text through explicit phrase-membership policy. It is not taint analysis, prompt-injection defense, sandboxing, information-flow control, or launch-scope decision.
Research & science (9)
Finance Forecast Evaluation SpineRuns econometric forecast-evaluation tests on synthetic fixtures, recording p-values and refusals with no advice.4/5Runs real tools
Job Runs econometric forecast-evaluation tests on synthetic fixtures, recording p-values and refusals with no advice.
Scope limit synthetic fixture forecast-evaluation statistics only; no investment-related actions, live market data, track record, or performance claim
Prediction Market Board BundleReplays imported quant market math on test rows, with duplicate retention and seven refusals.5/5
Job Replays imported quant market math on test rows, with duplicate retention and seven refusals.
Scope limit This is deterministic fixture evidence for copied quant helpers only; it is not live prediction-market-level conclusions, not provider truth, not forecast correctness, not investment-related actions, not external model access, and not launch-scope decision.
Market Dashboard Read-Model BundleRuns a copied market-dashboard reader to catch broken links, stale feeds, and trading overclaims.5/5
Job Runs a copied market-dashboard reader to catch broken links, stale feeds, and trading overclaims.
Scope limit This is fixture-bound read-model, freshness, and relation-grouping evidence only; it is not live market-level conclusions, not investment-related actions, not external model access, not launch-scope decision, and not whole-system correctness.
Toy-Transformer Attribution ReplayRecords which model features drove an answer, each tied to checkable evidence.4/5
Job Records which model features drove an answer, each tied to checkable evidence.
Scope limit It validates only the declared public circuit-attribution runtime-result record contract. It excludes model-transparency product claims, live model access, export of private weights/raw activations/proprietary prompts/hidden chain-of-thought, external model access, benchmark claims, or public sharing/launch.
Gridworld Counterfactual State ReplayReplays six what-if robotics scenes to show what a spatial prediction claim is built from.4/5
Job Replays six what-if robotics scenes to show what a spatial prediction claim is built from.
Scope limit It validates only the declared public contract of synthetic spatial counterfactual-replay metadata rows. It is evidence for inspectable replay rows and limitation labels, not for real-world spatial accuracy, simulator-product validity, media-only authority, operational deployment, service distribution, or scope decisions.
Research Replication Rubric Artifact ReplayAudits whether a paper-replication claim carries the full evidence trail.3/5
Job Audits whether a paper-replication claim carries the full evidence trail.
Scope limit It validates the shape and presence of synthetic replay metadata and result record references only - it does not run any experiment, metric script, or rerun, excludes any claim that a paper was actually replicated, that a benchmark claims was achieved, or that the underlying science is correct, and it never calls providers, exposes private paper/data bodies, or authorizes public sharing or launch.
Materials Lab-Safety Refusal ReplayReplays a self-driving lab loop as records, with safety gates and no real chemicals, robot, or lab.3/5
Job Replays a self-driving lab loop as records, with safety gates and no real chemicals, robot, or lab.
Scope limit It documents projection and replay mechanics only and excludes wetlab protocols, hazardous synthesis steps, reagent amounts, controlled/bioactive targets, robot commands, live assay data, discovery claims, benchmark claims, external model access, or any judgment of domain/chemical correctness.
Prediction Oracle ReconciliationReplays a forecast against the discipline a careful predictor would have to defend.3/5
Job Replays a forecast against the discipline a careful predictor would have to defend.
Scope limit It exercises projection mechanics on a synthetic, invented packet only. It does not establish forecasting correctness or accuracy, give trading/financial/investment-related actions, call live market data or providers, publish predictions, claim any performance or track record, import non-public data, or include launch operations.
Derived Fact Provider RuntimeFills in facts from simple recipes and records a clear error when one points at missing data.4/5Runs real tools
Job Fills in facts from simple recipes and records a clear error when one points at missing data.
Scope limit The bundle demonstrates registry-backed JSON-pointer, glob-count, and git-backed callable fact providers with provider failures represented as error rows. It is not a doctrine truth auditor, not a full source registry export, not semantic claim validation, and not launch-scope decision.
Import & drift control (20)
Source Projection Import ProtocolGates private-to-public imports, accepting only files with matching fingerprints and sources.5/5
Job Gates private-to-public imports, accepting only files with matching fingerprints and sources.
Scope limit It authorizes only verified source body import with provenance and content-digest checks; it does not grant source authority, whole-system equivalence, launch, hosted deployment, public sharing, recipient work, provider or Lean/Lake execution, secret or private-source-body export, or any whole-system correctness claim.
Projection-Drift Contract ValidatorPinpoints where a projected world-model copy drifted from its real source, with repair routes.5/5
Job Pinpoints where a projected world-model copy drifted from its real source, with repair routes.
Scope limit It only validates the declared public, metadata-only drift-result record contract. It supports inspection of recorded drift rows and source-linked refs; live repair, source control, doctrine changes, model-output export, public sharing, and launch are outside the fixture. It does not claim complete drift coverage or live repair control.
Unsurfaced Source Primitives BundleExposes eleven real but under-surfaced parts and rejects non-public-state and overclaim cases.5/5
Job Exposes eleven real but under-surfaced parts and rejects non-public-state and overclaim cases.
Scope limit It validates only a public source-open bundle and bounded public exercises; it is not raw operator memory, not prompt-shelf capture authority, not live market data, not provider/browser state, not media launch, and not public sharing or launch-scope decision.
Authority Systems Source BundleReplays eight authority and systems checks, rejecting provider, proof, and launch overclaims.5/5
Job Replays eight authority and systems checks, rejecting provider, proof, and launch overclaims.
Scope limit It validates only copied Set 5 authority-system source bodies and bounded deterministic exercises; it does not dispatch providers, prove Lean success, send live process signals, mutate generated state, change source files, authorize public sharing, include launch operations, or claim whole-system equivalence.
Trace, Code-Map & Scheduling Engines BundleRuns fifteen trace, code-map, and scheduling engines on test data, blocking truth overclaims.5/5
Job Runs fifteen trace, code-map, and scheduling engines on test data, blocking truth overclaims.
Scope limit It validates only a public source-open bundle and bounded exercises; it is not live source authority, whole-system equivalence, semantic truth, investment-related actions, complete sandbox proof, selected-test sufficiency proof, public sharing, or launch-scope decision.
Oracle Sibling Source BundleReplays subject-index and truth-diff logic on copied code, rejecting reasoning overclaims.5/5
Job Replays subject-index and truth-diff logic on copied code, rejecting reasoning overclaims.
Scope limit It validates only Oracle sibling copied source bodies and bounded deterministic exercises; it does not run Oracle reasoning, dispatch providers or bridges, invoke private orchestration engine, change source files, prove semantic truth, prove all Oracle paths are covered, authorize public sharing, or include launch operations.
Demo Take Console Source BundleReplays the recording console's Swift logic without launching the app or capturing audio.5/5
Job Replays the recording console's Swift logic without launching the app or capturing audio.
Scope limit It validates only Demo Take Console copied Swift source bodies and bounded deterministic exercises; it does not launch the app, authorize screen or microphone capture, export recording sessions, execute FFmpeg, dispatch WhisperKit or other models, change source files, prove complete UI coverage, authorize public sharing, or include launch operations.
Tools-Tail Primitives BundleExercises four copied helper tools over fixed inputs without touching live systems or data.5/5
Job Exercises four copied helper tools over fixed inputs without touching live systems or data.
Scope limit It validates only the imported source body. It does not claim source authority, whole-system equivalence, launch, or public sharing.
Policy Engines BundleMaps three policy engines over test data without model calls or live campaign execution.5/5
Job Maps three policy engines over test data without model calls or live campaign execution.
Scope limit It validates only the imported source body. It does not claim source authority, whole-system equivalence, launch, or public sharing.
Audio Level RMS PortComputes the audio loudness math on test arrays without opening a microphone or capturing input.3/5
Job Computes the audio loudness math on test arrays without opening a microphone or capturing input.
Scope limit projection mechanics only, not domain-level conclusions
Structural Theses Finance BundleRuns a copied finance-thesis model through dated test cases with no live market data or advice.5/5
Job Runs a copied finance-thesis model through dated test cases with no live market data or advice.
Scope limit It validates only the imported source body over synthetic thesis rows. It does not claim source authority, whole-system equivalence, financial or investment-related actions, live market data, portfolio action, external model access, launch, public sharing, launch, or public sharing.
Engine Room DemoRuns proof, runtime, security, and routing demos through bounded public examples with stated limits.5/5
Job Runs proof, runtime, security, and routing demos through bounded public examples with stated limits.
Scope limit It validates only the public Engine Room composition contract; it is not deployment posture, whole-system equivalence, frontier theorem proving, complete security proof, public sharing, or launch-scope decision.
Backend & Governance Engines BundleExercises thirteen copied backend and governance engines over fixed public test cases.5/5
Job Exercises thirteen copied backend and governance engines over fixed public test cases.
Scope limit It validates only a public source-open bundle and bounded synthetic exercises; it is not live lineage truth, human approval authority, market/news truth, host-state truth, work log truth, external model access, source-file changes, public sharing, launch-scope decision, or whole-system equivalence.
Governance & Compiler Mechanisms BundleChecks thirteen copied governance and compiler routines against the code they were copied from.5/5
Job Checks thirteen copied governance and compiler routines against the code they were copied from.
Scope limit It validates only the imported source body. It does not claim source authority, whole-system equivalence, launch, public sharing, live ledger control, or source-file changes.
Saturation Engines BundleVerifies twelve copied engine routines and computes each failure probe from inputs, not echoes.5/5
Tool Server Pressure InventoryFlags detached helper processes and launch pressure from synthetic rows, not live hosts.5/5
Job Flags detached helper processes and launch pressure from synthetic rows, not live hosts.
Scope limit validates declared public helper-process pressure inventory contract only; no live process reads, process signalling, host mutation, launch-scope decision, external model access, non-public data equivalence, or whole-system correctness
Compliance Pipeline BundleConfirms six copied compliance source files carry their functions; runs one helper on sample text.3/5
Job Confirms six copied compliance source files carry their functions; runs one helper on sample text.
Scope limit validates declared public Set 8 compliance pipeline bundle contract only; no full compliance-ledger freshness, external model access, model dispatch, source-file changes, source note mutation, launch, public sharing, non-public data equivalence, or whole-system correctness
Live Source Drift BundleCompares four copied router and landing routines against current code to surface stale copies.5/5
Job Compares four copied router and landing routines against current code to surface stale copies.
Scope limit verified source body import only, not route authority, work log or work log mutation authority, mission-transaction execution, git staging or commit approval, source-file changes, non-public runtime export, launch, or public sharing
Release Public Wording GateFlags affirmative open-source and deployment-posture wording while allowing safe boundary notes.5/5
Job Flags affirmative open-source and deployment-posture wording while allowing safe boundary notes.
Scope limit This is lexical fixture evidence only; it is not launch-scope decision, not publishing-scope decision, not semantic NLP truth, not secret-scan coverage, and not whole-system correctness.
Generated Projection Drift RuntimeRuns each owner's own drift check on generated files and flags repairs, without changing the file.4/5Runs real tools
Job Runs each owner's own drift check on generated files and flags repairs, without changing the file.
Scope limit The bundle demonstrates owner-routed generated projection drift detection over declared artifacts, source authorities, clean-result record fingerprints, and no-write check return codes. It is not semantic drift proof, not full source registry validation, not repair authority, and not launch-scope decision.
Work & continuity (5)
Mission Transaction Work SpineRuns the real work-ledger engine on a sanitised snapshot to re-derive each change's verdict.4/5Runs real tools
Job Runs the real work-ledger engine on a sanitised snapshot to re-derive each change's verdict.
Scope limit It validates work-landing, claim, checkpoint-lane, and dependency metadata projections over fixed fixtures only; it does not mutate live ledgers or git, certify real completion, authorize broad staging without operator intent, or prove any change is actually correct or complete.
Durable Agent Work Landing ReplayAudits recorded work-claims so each cites files, validates before commit, and proves HEAD moved.5/5
Job Audits recorded work-claims so each cites files, validates before commit, and proves HEAD moved.
Scope limit It validates only the declared public work-landing contract over recorded rows. It is evidence for fixture-local completion mechanics, not for live Git side effects, unrelated-path staging, non-public body export, service operation, or distribution clearance.
Bridge-Continuity Sign-off ReplayReplays a paused job to prove the rules for safely resuming it hold and reject duplicate resumes.5/5
Job Replays a paused job to prove the rules for safely resuming it hold and reject duplicate resumes.
Scope limit It validates only the declared public continuity contract over synthetic fixtures; it does not run live bridge transport, use external model services, read operator HUD/browser/phase-runtime or private-memory state, prove provider or UI uptime, land work, change source files, or include launch operations.
Concurrency Mission ControlRuns copied claim-coordination code so duplicate, stale, and conflicting claims get blocked.5/5
Job Runs copied claim-coordination code so duplicate, stale, and conflicting claims get blocked.
Scope limit verified concurrency mission-control source body import only, not a live scheduler, external model access, hosted orchestration, production concurrency-safety proof, source authority, whole-system equivalence, public sharing, or launch-scope decision
Semantic Singleflight Dedup RuntimeReuses one result for duplicate command runs, keyed by repo state rather than the words alone.4/5Runs real tools
Job Reuses one result for duplicate command runs, keyed by repo state rather than the words alone.
Scope limit It keys and dedups command runs by a repo-state fingerprint over bounded public fixture commands only; it does not guarantee global mutual exclusion, does not replace a lock service, cannot prove cross-host correctness, and is not a job scheduler, a daemon, or launch-scope decision.