Plectis
This page

Reference

Components

A generated index over the 88 public components, grouped by the seven source families. Each card is a terse spec and evidence record: what it does, its scope limit, and how strong the evidence is. For the long-form reasoning behind a component, follow its paper module. Every field is projected directly from source.

Entry & orientation (2)

Cold Reader Route MapVerifies the first-run guided path so every step names a real command, doc, and evidence.5/5

Job Verifies the first-run guided path so every step names a real command, doc, and evidence.

Scope limit It is projection-only metadata that validates the declared public route contract; it is not route registry control and excludes source-file changes, external model access, launch/public sharing, financial decisions, non-public data equivalence, or whole-system correctness.

Public Reveal WalkthroughBinds the first-time reader tour to evidence so each count leads to a source.4/5Runs real tools

Job Binds the first-time reader tour to evidence so each count leads to a source.

Scope limit It authorizes only bounded public reveal runtime behavior and a digest-verified public body-import witness; it excludes launch, hosted deployment, public sharing, recipient work, external model access, secret export, non-public data equivalence, Lean/Lake execution, whole-system correctness, or general product authority.

EvidenceBounded runtime computationevidence 4/5Real runtime result

getting-startedinteresting-partsevaluation

Architecture & navigation (12)

Pattern Binding ContractChecks a real pattern catalog for digest, cross-reference, and dependency-cycle integrity.5/5

Job Checks a real pattern catalog for digest, cross-reference, and dependency-cycle integrity.

Scope limit It validates only the declared public pattern-binding/route-readiness contract; it does not certify the private pattern ledger, public launch or hosted-public posture, public sharing, external model access, non-public data equivalence, or whole-system correctness, and it does not turn any mined pattern row into a standalone public leaf (selection stays component-first and fixture-bound).

Pattern Assimilation StepVerifies each landed task filed exactly one learning record naming what it changed.5/5

Job Verifies each landed task filed exactly one learning record naming what it changed.

Scope limit It validates only the declared public completion contract over synthetic fixture data; it does not ingest private lessons, mutate live ledgers, promote global doctrine, include launch operations or public sharing, make external model access, claim non-public data equivalence, or certify public runtime behavior.

Executable Doctrine GrammarChecks that example standards files declare their purpose, rule, records, and what they do not claim.5/5
Navigation Hologram Route PlaneAudits a folder's navigation so browse rows never pose as the source of truth.5/5

Job Audits a folder's navigation so browse rows never pose as the source of truth.

Scope limit It validates only the declared public toy route-plane contract and its regression fixtures (plus exact copied navigation source modules in the bundle path); it does not establish live route freshness, grant source authority, authorize any later component, run any provider/live-kernel call, or certify the whole wave.

Standards Meta DiagnosticsConfirms every accepted part still ties to a written rule, a run command, and a saved proof.5/5
Voice To Doctrine Self Improvement LoopVerifies each lesson changed a named owner page with evidence before the loop closes.5/5

Job Verifies each lesson changed a named owner page with evidence before the loop closes.

Scope limit It validates only the declared contract of the loop on fixtures; it does not export source notes or private bodies, grant source/doctrine edits, global-promotion, live work log mutation, or publishing-scope decision, make external model access, prove correctness, or claim whole-system equivalence.

Cognitive Operator RegistryChecks the catalog of named thinking-moves so each is fully described and backed by evidence.5/5
Routing Anti Patterns RegistryIndexes the navigation mistakes agents repeat and guards the public list.5/5
Doctrine Fact Claim AuditChecks that public fact rows state the right count and point at live, anchored code.5/5

Job Checks that public fact rows state the right count and point at live, anchored code.

Scope limit fact assertion, code-loci, and DAG fixture truth gate only; it is not a comprehension engine and does not establish a minimum read graph

Self Ignorance Coverage LedgerCompares expected against built entities to report known coverage gaps.3/5

Job Compares expected against built entities to report known coverage gaps.

Scope limit known Kind Atlas coverage debt projection only; it does not claim literal unknown-unknown omniscience or absence proof

Navigation Fitness BenchmarkMeasures a navigation result against its target for recall, precision, and speed on public examples.4/5Runs real tools

Job Measures a navigation result against its target for recall, precision, and speed on public examples.

Scope limit The bundle demonstrates a route-packet evaluator for expected stable ids, forbidden first routes, and latency budgets. It is not a live private kernel run, not an embedding benchmark, not a universal navigation benchmark, and not launch-scope decision.

Reference Knowledge RoutingRuns a keyword-overlap match that ranks catalog entries, shows why each matched, and admits when none do.4/5Runs real tools

Job Runs a keyword-overlap match that ranks catalog entries, shows why each matched, and admits when none do.

Scope limit The bundle demonstrates explainable tiered weighted-token retrieval over a sanitized reference catalog. It is not BM25, TF-IDF, embedding search, repository cloning, license authority, private reference corpus authority, or launch-scope decision.

Formal math & proof (20)

Certificate Kernel Execution LabRuns the Lean verifier over a small public proof project and reports which rows it accepted.4/5Runs real tools

Job Runs the Lean verifier over a small public proof project and reports which rows it accepted.

Scope limit It is a local tool-witness that the declared public fixture rows compiled and were adjudicated by the local Lean verifier; it excludes general proof authority, count oracle/provider output as proof, expose proof text, change source files, claim a benchmark solve-rate, or include launch operations.

Formal Math Lean Proof WitnessCompiles a tiny Lean example with the real prover and records whether it built, leaking no proof text.4/5Runs real tools

Job Compiles a tiny Lean example with the real prover and records whether it built, leaking no proof text.

Scope limit It authorizes only a witness that a tiny declared public toy proof compiled under the locally installed Lean/Lake toolchain in a temporary workspace, plus confirmation that its leakage guardrails fired. It excludes Mathlib/Aesop/Batteries-dependent or general proof or theorem-program authority, external model access, private proof import, benchmark or performance claims, whole-system correctness, or any launch, hosted deployment, or public sharing.

Verifier Lab Execution SpineRuns Lean on small bounded proof attempts in a temp copy and records what passed or failed.4/5Runs real tools

Job Runs Lean on small bounded proof attempts in a temp copy and records what passed or failed.

Scope limit It is a tool-witness result record for bounded public Lean transition rows only: it does not establish general proof authority, count oracle/provider output as proof, export proof bodies or tactic scripts, use external model services, change source files, claim benchmark solve-rates, or include launch operations/public sharing.

Corpus Readiness Mathlib Absence GateRuns the real Lean toolchain to confirm the math library is absent, then gates proof tasks.4/5Runs real tools

Job Runs the real Lean toolchain to confirm the math library is absent, then gates proof tasks.

Scope limit It only projects and gate-checks recorded corpus/toolchain readiness accounting, re-verifies recorded source digests and leakage guards, and runs a bounded Lean/Lake import probe when a toolchain is present. It does not run a full Lake build, prove formal-result correctness, claim Mathlib is available beyond the probe result, benchmark corpora, score model performance, use external model services, or include launch operations or public sharing.

Proof / Control / Runtime Import BundleChecks fourteen proof, control, and runtime parts as one unit that rejects every overclaim.5/5
Proof Diagnostic Evidence SpineSorts proof-pipeline checks into accepted or rejected without inflating a pass.3/5

Job Sorts proof-pipeline checks into accepted or rejected without inflating a pass.

Scope limit It records proof/evidence diagnostics over existing result record references only. It does not run Lean, use external model services, expose proof bodies, turn a passing check into formal-proof or theorem authority, prove runtime or whole-system correctness, authorize later components, certify public launch, authorize public sharing or recipient work, or establish secret export.

Formal Math Readiness GateReads declared math setups and lists which proof tactics may be attempted versus blocked.3/5

Job Reads declared math setups and lists which proof tactics may be attempted versus blocked.

Scope limit It only validates and projects declared readiness metadata; it does not run Lean/Lake, inspect the real toolchain, use external model services, prove any theorem correct, produce benchmark claims, or authorize Mathlib-dependent proof attempts.

Mathematical Strategy Atlas Hypothesis ScorerPicks a first-guess proof strategy from a problem's tags and flags any it cannot map.3/5

Job Picks a first-guess proof strategy from a problem's tags and flags any it cannot map.

Scope limit It only projects pre-oracle strategy-hypothesis and retrieval mechanics; it does not run Lean/Lake, prove theorems, establish domain or formal-result correctness, reveal oracle labels, expose proof bodies, use external model services, tune on test answers, or include launch operations.

Tactic Portfolio Availability ProbeMaps which Lean proof tactics a recorded run marked usable before any code relies on one.3/5

Job Maps which Lean proof tactics a recorded run marked usable before any code relies on one.

Scope limit It only projects and validates which tactics were recorded as compiling in one captured environment; it does not run Lean/Lake at all, prove any goal, certify domain-level conclusions, use external model services, claim benchmark performance, or include launch operations.

Target Shape Tactic Routing GateRecords an allow-or-reject decision and reason for each proof tactic before any proof runs.3/5

Job Records an allow-or-reject decision and reason for each proof tactic before any proof runs.

Scope limit It only inspects and records the projection mechanics of pre-execution tactic-routing referencesemitting per-tactic allow/reject decisions with reasons. It does not run Lean/Lake, does not establish or judge the correctness of any goal, emits no proof bodies, makes no external model access, performs no post-execution route selection, reports no benchmark claims or maturity, and excludes launch.

Lean Std Premise IndexLists a fixed catalog of public Lean building blocks and confirms none hides proof text or test answers.3/5

Job Lists a fixed catalog of public Lean building blocks and confirms none hides proof text or test answers.

Scope limit It only validates the projection of premise metadata and copied source bodies; it does not run Lean or Lake, prove any theorem correct, expose proof bodies or oracle-needed ids, use external model services, produce benchmark claims, or include launch operations.

Formal Math Premise RetrievalShows which lemmas a plain search surfaces per query, and never leaks proof text or answer keys.3/5

Job Shows which lemmas a plain search surfaces per query, and never leaks proof text or answer keys.

Scope limit It only checks that public retrieval metadata is internally coherent, term-scored over a copied index, budget-bounded, and leakage-clean; it does not run Lean/Lake, use external model services, prove any theorem or its own correctness, claim benchmark performance, or include launch operations.

Formal Math Verifier Trace Repair LoopReplays how a proof lab turns verifier failures into fixes, with no promotion without a fresh re-run.3/5

Job Replays how a proof lab turns verifier failures into fixes, with no promotion without a fresh re-run.

Scope limit It demonstrates control-loop projection mechanics over copied Ring2 run rows only; it does not run Lean/Lake, use external model services, expose proof bodies or oracle premise ids, treat human or provider advice as correctness, prove any theorem, or include launch operations.

Formal Evidence Cell Anchor ResolverResolves each proof-flavored math claim to named evidence and flags ones that overreach or lack backing.3/5

Job Resolves each proof-flavored math claim to named evidence and flags ones that overreach or lack backing.

Scope limit It validates claim-to-evidence anchoring mechanics only: claim-to-cell resolution, source-anchor presence, permitted claim strength, copied-source-module digest checks, and leakage refusals. It does not run Lean/Lake, certify theorem or mathematical correctness, expose proof bodies or non-public source refs, use external model services, or include launch operations/public sharing.

Undeclared Library Prior Symbol ClassifierDetects when a checked Lean proof cites a library result outside its approved set.3/5

Job Detects when a checked Lean proof cites a library result outside its approved set.

Scope limit It only projects the symbol-boundary classification mechanic over copied Lean/Std premise rows and pre-extracted symbol observations; it does not read proof source, run Lean or Lake, prove formal-result correctness, treat the whole standard library as an implicit allowlist, claim Mathlib availability, use external model services, or include launch operations.

Ring2 Premise Retrieval Precision Recall HarnessScores how much proof support a premise search found, problem by problem.3/5

Job Scores how much proof support a premise search found, problem by problem.

Scope limit These are after-the-fact retrieval-attribution labels and precision/recall counts over copied run records only. The component does not run Lean or Lake, call any provider, expose proof bodies, tune on test answers, claim benchmark performance, prove formal-result correctness, or include launch operations, and its labels are explicitly forbidden from flowing into provider context. The aggregate numbers describe only the copied fixture/bundle replayed, not any benchmark claims.

Verifier Lab KernelFolds nine proof checks into one report labeling each line by which source actually backs it.5/5

Job Folds nine proof checks into one report labeling each line by which source actually backs it.

Scope limit It validates the declared public contract shape of the proof packet and component result records only; it does not establish anything correct, count oracle/provider output as forward proof success, import private or Mathlib-dependent proof bodies, use external model services, change source files, or claim benchmark solve rates, launch, or maturity.

Proof Derived Governed Mutation AuthorizationChecks a synthetic change-authorization record for its proof-and-approval chain, bound to a real commit.5/5
Finite Erdos Denominator-Order Certificate StrikeComputes an exact-arithmetic denominator-order identity and catches forged ones, not the open Erdos problem.4/5Runs real tools

Job Computes an exact-arithmetic denominator-order identity and catches forged ones, not the open Erdos problem.

Scope limit It computes the finite denominator-order certificate ord_Q(b)=lcm(F) for S_F(b)=sum 1/(b^n-1)=P/Q in exact rational arithmetic over bounded public fixtures and rejects forged certificates by recomputation; it does not establish the open infinite Erdos #257 problem, is not an oracle, prover, or provider result, and a holding certificate is a bounded computational witness, not a machine-checked proof.

Lean Proof-Search Lab RuntimeFinds tactic scripts the real Lean prover accepts on toy theorems, and refuses proofs that cheat.4/5Runs real tools

Job Finds tactic scripts the real Lean prover accepts on toy theorems, and refuses proofs that cheat.

Scope limit This standard governs a bounded symbolic Lean proof-search lab over tiny public toy-theorem fixtures whose verdict source is the installed Lean subprocess only. It is not neural theorem proving, does not solve any open mathematical problem, does not forward oracle proof bodies, is not frontier-scale math automation or online-RL bandit search, is not a private source prover-run export, and excludes launch or public sharing.

Agent reliability & safety (20)

Agent Completion Faithfulness AuditRuns real git and pytest on a sample repo so wrap-up claims state only what the evidence proves.4/5Runs real tools

Job Runs real git and pytest on a sample repo so wrap-up claims state only what the evidence proves.

Scope limit verified means the referenced evidence object exists or a pytest span ran; it does not imply the span passed unless exit-zero status was explicitly checked

EvidenceExternal tool runevidence 4/5Real runtime result

ai-safetyagent-evaluationred-teaming

Bounded Autonomy Campaign PacketDrafts proposed work from coverage gaps and proves it cannot repair or rewrite the code itself.4/5Runs real tools

Job Drafts proposed work from coverage gaps and proves it cannot repair or rewrite the code itself.

Scope limit self-proposal campaign packet only; no self-repair or unsupervised source-file changes

EvidenceExternal tool runevidence 4/5Real runtime result

ai-safetyagent-evaluationred-teaming

Provider Context Recipe Budget PolicyRuns the real context harness to measure assembled byte sizes and check each bundle fits its budget.4/5Runs real tools

Job Runs the real context harness to measure assembled byte sizes and check each bundle fits its budget.

Scope limit It validates context-budget projection mechanics (byte ceilings, ordered section fill, omitted-section manifests, deliverable routing, and digest-checked source-body imports) only. It excludes provider/API calls, run Lean/Lake, expose or carry proof or oracle truth-side material, assert theorem or domain-level conclusions, or include launch operations.

EvidenceBounded runtime computationevidence 4/5Real runtime result

ai-safetyagent-evaluationred-teaming

Cold Evaluation Honesty BundleRuns a copied route-quality simulator and checks its all-B scorecard against the original code.5/5

Job Runs a copied route-quality simulator and checks its all-B scorecard against the original code.

Scope limit verified cold-eval source body import only, not a live benchmark, navigation truth, source authority, external model access, whole-system equivalence, public sharing, or launch-scope decision

EvidenceVerified source importevidence 5/5Copied source body

ai-safetyagent-evaluationred-teaming

Validator Checker BundleRuns the real validator code over public examples so its safety checks stay inspectable.5/5
Secondary Runtime Source BundleRuns eight trace, graph, and market engines on test rows without fetching live markets.5/5

Job Runs eight trace, graph, and market engines on test rows without fetching live markets.

Scope limit verified source body import only; no browser/session export, wallet authority, live market data, investment-related actions, external model access, source-file changes, whole-system equivalence, public sharing, launch, semantic-truth, or whole-system correctness claim

Agent Benchmark Integrity Anti Gaming ReplayValidates a synthetic benchmark-integrity record and flags the contamination cases it declares.3/5

Job Validates a synthetic benchmark-integrity record and flags the contamination cases it declares.

Scope limit It authorizes only bounded public runtime validation over copied source-open pattern provenance bodies and metadata-only benchmark-integrity replay rows; it does not establish any benchmark or SWE-bench score, agent capability, external model service, live-repo mutation, private/oracle/hidden-gold body access, product progress, or launch-scope decision.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

ai-safetyagent-evaluationred-teaming

Monitor Evidence-Boundary ReplayReplays honest and deceptive agent runs and flags any verdict missing its declared backing evidence.3/5

Job Replays honest and deceptive agent runs and flags any verdict missing its declared backing evidence.

Scope limit Bounded public runtime validation over copied source pattern bodies, sanitized dogfood trace slices, recomputed monitor-verdict spans, source-artifact evidence refs, digest/metadata-only/non-public-state gates, and negative cases only; no live agent execution, monitor product performance, control-eval score, safety-validation, benchmark, provider-call, source-file changes, launch, public sharing, or product authority.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

ai-safetyagent-evaluationred-teaming

Sabotage-Monitor Contract ReplayAudits a hidden-goal catch claim for the steps, suspicion scores, and counterfactual it needs.3/5

Job Audits a hidden-goal catch claim for the steps, suspicion scores, and counterfactual it needs.

Scope limit Bounded public runtime validation over copied source pattern bodies, sanitized dogfood trace slices, recomputed sabotage/scheming monitor spans, source-artifact evidence refs, digest/metadata-only/non-public-state gates, and negative cases only; no live sabotage, live agent execution, exploit instruction, account secret/account, private-reasoning, harmful-payload, monitor-product-performance, deployment-risk, benchmark, provider-call, source-file changes, launch, public sharing, or product authority.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

ai-safetyagent-evaluationred-teaming

Agent Memory Temporal Conflict ReplayReplays a memory edit-and-delete to show stale facts get flagged before they sway an answer.3/5

Job Replays a memory edit-and-delete to show stale facts get flagged before they sway an answer.

Scope limit It validates the projection mechanics of a synthetic memory fixture only — that the required refs, decisions, paired replays, negative cases, and secret-exclusion scan line up and that result records are metadata-only. It does not claim live-memory product quality, judge whether memory decisions were domain-correct, treat memory recall as source authority, adopt active injection, export private transcripts, use external model services, change source files, or include launch operations.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

ai-safetyagent-evaluationred-teaming

Memory-Poisoning Quarantine Policy ReplayReplays a recorded memory-tamper case, checking its declared quarantine, block, and delete steps line up.3/5

Job Replays a recorded memory-tamper case, checking its declared quarantine, block, and delete steps line up.

Scope limit It only checks the structural shape and internal consistency of a synthetic memory-security policy projection recorded as JSON. It does not run or validate any real memory store, does not itself quarantine, delete, or re-run anything, and does not establish that any system actually resists poisoning. It exports no private memory bodies or transcripts, calls no providers, mutates no source, produces no benchmark claims, and excludes launch (all scope limit flags are hardcoded false).

EvidenceComputed projectionevidence 3/5Source-faithful refactor

ai-safetyagent-evaluationred-teaming

MCP Tool-Authority Policy ReplayAudits a recorded tool-use log to confirm each action was scoped, approved, undoable, and fenced.3/5

Job Audits a recorded tool-use log to confirm each action was scoped, approved, undoable, and fenced.

Scope limit It only checks that the tool-authority evidence in a recorded bundle (scopes, approvals, rollbacks, instruction/data splits, cold replays, redaction, and the expected abuse-case failures) is present and internally consistent. It does not run tools or authorize live MCP/account access, account secret or payload export, treating tool output as instruction, source-file changes, benchmark safety scores, or launch, and it makes no claim that the underlying tool-use policy is domain-correct.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

ai-safetyagent-evaluationred-teaming

Belief-State Reward Bundle ReplayChecks that each step reward in a recorded run cites a declared verifier-feedback row, not a trick.3/5

Job Checks that each step reward in a recorded run cites a declared verifier-feedback row, not a trick.

Scope limit It only checks that the projection's accounting lines up under its own schema rules over recorded synthetic fixtures; it excludes hidden-reasoning export, RL training, hidden gold or neural-judge-only labels, benchmark-performance claims, external model access, source-file changes, or launch, and proves nothing about real-world reward, live agent behavior, or domain-level conclusions.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

ai-safetyagent-evaluationred-teaming

Sandbox-Policy ReplayMaps sandboxed agent actions to show each was approved or blocked before running, then rolled back.3/5

Job Maps sandboxed agent actions to show each was approved or blocked before running, then rolled back.

Scope limit It validates the projection / trace-refactor mechanics over a synthetic fixture only; it excludes live sandbox escape, secret or account secret handling, live network access, host filesystem mutation, executable payload export, raw environment export, external model access, security benchmark claims, source-file changes, or launch. A pass proves the projection boundary and trace-refactor mechanics for this contract, not real sandbox security, exploit resistance, or whole-system safety.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

ai-safetyagent-evaluationred-teaming

Prompt-Injection Flow-Policy ReplayReplays an agent run to show untrusted text was gated before any sensitive action, leaking no secret.3/5

Job Replays an agent run to show untrusted text was gated before any sensitive action, leaking no secret.

Scope limit Passing result records only show this projection satisfies the named information-flow contract over synthetic, redacted, metadata-only rows; they do not prove general prompt-injection robustness, benchmark performance, live account/tool/provider safety, hidden-message handling in a real system, source-file changes, or launch-scope decision.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

ai-safetyagent-evaluationred-teaming

Vulnerability Patch-Proof ReplayChecks a fixed-bug evidence chain and re-runs three small real security checks; no real attack material.3/5

Job Checks a fixed-bug evidence chain and re-runs three small real security checks; no real attack material.

Scope limit It validates only the projection/evidence-chain mechanics of a synthetic replay: structural presence, cross-reference consistency, declared boolean flags, and the secret/live-access exclusion scan. It executes small regression witnesses but performs no real vulnerability discovery and makes no judgment of real-world security or fix correctness. It excludes live-target testing, real CVE exploitation, weaponized payloads, account secret handling, network exfiltration, actionable exploit steps, external model access, source-file changes, benchmark security scores, launch, or any whole-system security claim.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

ai-safetyagent-evaluationred-teaming

Agent Route Observability RuntimeRecomputes an agent run's route-compliance score and anti-pattern flags with real trace-analytics code.5/5

Job Recomputes an agent run's route-compliance score and anti-pattern flags with real trace-analytics code.

Scope limit It validates only public, recorded trace-feedback metadata and regression fixtures; it does not inspect live operator state, certify or prove runtime behavior, read model-output data, mutate the work log, authorize pattern assimilation, or include launch operations.

Bridge Campaign DAG ValidationRuns shape checks on a fan-out work plan for unique steps, dependencies, and no cycles, not the plan itself.4/5Runs real tools

Job Runs shape checks on a fan-out work plan for unique steps, dependencies, and no cycles, not the plan itself.

Scope limit The bundle validates a public bridge-campaign DAG contract and provider worker ceiling. It is not a dispatcher, not a live multi-agent run, not a provider safety proof, and not launch-scope decision.

Metabolism Queue ReconciliationRuns a scratch-database model of a durable job queue and flags impossible states for a human to review.4/5Runs real tools

Job Runs a scratch-database model of a durable job queue and flags impossible states for a human to review.

Scope limit The bundle demonstrates a synthetic SQLite durable queue, lease recovery, blackboard claim-event projection, and cold-start reconciliation taxonomy. It is not a live non-public runtime export, not an agent dispatcher, not external model service, not ambiguous auto-repair, not a distributed database, and not launch-scope decision.

EvidenceBounded runtime computationevidence 4/5Real runtime result

agent-concurrencycontinuityoperational-discipline

Egress Self-Compliance AuditRuns phrase-membership checks on an agent's own replies for three self-policing slips, not deep meaning.4/5Runs real tools

Job Runs phrase-membership checks on an agent's own replies for three self-policing slips, not deep meaning.

Scope limit The bundle evaluates egress text through explicit phrase-membership policy. It is not taint analysis, prompt-injection defense, sandboxing, information-flow control, or launch-scope decision.

EvidenceBounded runtime computationevidence 4/5Real runtime result

ai-safetycompliancered-teaming

Research & science (9)

Finance Forecast Evaluation SpineRuns econometric forecast-evaluation tests on synthetic fixtures, recording p-values and refusals with no advice.4/5Runs real tools

Job Runs econometric forecast-evaluation tests on synthetic fixtures, recording p-values and refusals with no advice.

Scope limit synthetic fixture forecast-evaluation statistics only; no investment-related actions, live market data, track record, or performance claim

EvidenceExternal tool runevidence 4/5Real runtime result

research-workflowsforecastingfinance

Prediction Market Board BundleReplays imported quant market math on test rows, with duplicate retention and seven refusals.5/5

Job Replays imported quant market math on test rows, with duplicate retention and seven refusals.

Scope limit This is deterministic fixture evidence for copied quant helpers only; it is not live prediction-market-level conclusions, not provider truth, not forecast correctness, not investment-related actions, not external model access, and not launch-scope decision.

EvidenceVerified source importevidence 5/5Copied source body

research-workflowsforecastingfinance

Market Dashboard Read-Model BundleRuns a copied market-dashboard reader to catch broken links, stale feeds, and trading overclaims.5/5

Job Runs a copied market-dashboard reader to catch broken links, stale feeds, and trading overclaims.

Scope limit This is fixture-bound read-model, freshness, and relation-grouping evidence only; it is not live market-level conclusions, not investment-related actions, not external model access, not launch-scope decision, and not whole-system correctness.

EvidenceVerified source importevidence 5/5Copied source body

research-workflowsforecastingfinance

Toy-Transformer Attribution ReplayRecords which model features drove an answer, each tied to checkable evidence.4/5

Job Records which model features drove an answer, each tied to checkable evidence.

Scope limit It validates only the declared public circuit-attribution runtime-result record contract. It excludes model-transparency product claims, live model access, export of private weights/raw activations/proprietary prompts/hidden chain-of-thought, external model access, benchmark claims, or public sharing/launch.

EvidenceContract validatorevidence 4/5Real runtime result

research-workflowsforecastingprovider operations

Gridworld Counterfactual State ReplayReplays six what-if robotics scenes to show what a spatial prediction claim is built from.4/5

Job Replays six what-if robotics scenes to show what a spatial prediction claim is built from.

Scope limit It validates only the declared public contract of synthetic spatial counterfactual-replay metadata rows. It is evidence for inspectable replay rows and limitation labels, not for real-world spatial accuracy, simulator-product validity, media-only authority, operational deployment, service distribution, or scope decisions.

Research Replication Rubric Artifact ReplayAudits whether a paper-replication claim carries the full evidence trail.3/5

Job Audits whether a paper-replication claim carries the full evidence trail.

Scope limit It validates the shape and presence of synthetic replay metadata and result record references only - it does not run any experiment, metric script, or rerun, excludes any claim that a paper was actually replicated, that a benchmark claims was achieved, or that the underlying science is correct, and it never calls providers, exposes private paper/data bodies, or authorizes public sharing or launch.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

research-workflowsforecastingprovider operations

Materials Lab-Safety Refusal ReplayReplays a self-driving lab loop as records, with safety gates and no real chemicals, robot, or lab.3/5

Job Replays a self-driving lab loop as records, with safety gates and no real chemicals, robot, or lab.

Scope limit It documents projection and replay mechanics only and excludes wetlab protocols, hazardous synthesis steps, reagent amounts, controlled/bioactive targets, robot commands, live assay data, discovery claims, benchmark claims, external model access, or any judgment of domain/chemical correctness.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

research-workflowsforecasting

Prediction Oracle ReconciliationReplays a forecast against the discipline a careful predictor would have to defend.3/5

Job Replays a forecast against the discipline a careful predictor would have to defend.

Scope limit It exercises projection mechanics on a synthetic, invented packet only. It does not establish forecasting correctness or accuracy, give trading/financial/investment-related actions, call live market data or providers, publish predictions, claim any performance or track record, import non-public data, or include launch operations.

EvidenceComputed projectionevidence 3/5Source-faithful refactor

research-workflowsforecastingprovider operations

Derived Fact Provider RuntimeFills in facts from simple recipes and records a clear error when one points at missing data.4/5Runs real tools

Job Fills in facts from simple recipes and records a clear error when one points at missing data.

Scope limit The bundle demonstrates registry-backed JSON-pointer, glob-count, and git-backed callable fact providers with provider failures represented as error rows. It is not a doctrine truth auditor, not a full source registry export, not semantic claim validation, and not launch-scope decision.

Import & drift control (20)

Source Projection Import ProtocolGates private-to-public imports, accepting only files with matching fingerprints and sources.5/5
Projection-Drift Contract ValidatorPinpoints where a projected world-model copy drifted from its real source, with repair routes.5/5

Job Pinpoints where a projected world-model copy drifted from its real source, with repair routes.

Scope limit It only validates the declared public, metadata-only drift-result record contract. It supports inspection of recorded drift rows and source-linked refs; live repair, source control, doctrine changes, model-output export, public sharing, and launch are outside the fixture. It does not claim complete drift coverage or live repair control.

Unsurfaced Source Primitives BundleExposes eleven real but under-surfaced parts and rejects non-public-state and overclaim cases.5/5

Job Exposes eleven real but under-surfaced parts and rejects non-public-state and overclaim cases.

Scope limit It validates only a public source-open bundle and bounded public exercises; it is not raw operator memory, not prompt-shelf capture authority, not live market data, not provider/browser state, not media launch, and not public sharing or launch-scope decision.

Authority Systems Source BundleReplays eight authority and systems checks, rejecting provider, proof, and launch overclaims.5/5

Job Replays eight authority and systems checks, rejecting provider, proof, and launch overclaims.

Scope limit It validates only copied Set 5 authority-system source bodies and bounded deterministic exercises; it does not dispatch providers, prove Lean success, send live process signals, mutate generated state, change source files, authorize public sharing, include launch operations, or claim whole-system equivalence.

Trace, Code-Map & Scheduling Engines BundleRuns fifteen trace, code-map, and scheduling engines on test data, blocking truth overclaims.5/5

Job Runs fifteen trace, code-map, and scheduling engines on test data, blocking truth overclaims.

Scope limit It validates only a public source-open bundle and bounded exercises; it is not live source authority, whole-system equivalence, semantic truth, investment-related actions, complete sandbox proof, selected-test sufficiency proof, public sharing, or launch-scope decision.

Oracle Sibling Source BundleReplays subject-index and truth-diff logic on copied code, rejecting reasoning overclaims.5/5

Job Replays subject-index and truth-diff logic on copied code, rejecting reasoning overclaims.

Scope limit It validates only Oracle sibling copied source bodies and bounded deterministic exercises; it does not run Oracle reasoning, dispatch providers or bridges, invoke private orchestration engine, change source files, prove semantic truth, prove all Oracle paths are covered, authorize public sharing, or include launch operations.

Demo Take Console Source BundleReplays the recording console's Swift logic without launching the app or capturing audio.5/5

Job Replays the recording console's Swift logic without launching the app or capturing audio.

Scope limit It validates only Demo Take Console copied Swift source bodies and bounded deterministic exercises; it does not launch the app, authorize screen or microphone capture, export recording sessions, execute FFmpeg, dispatch WhisperKit or other models, change source files, prove complete UI coverage, authorize public sharing, or include launch operations.

Tools-Tail Primitives BundleExercises four copied helper tools over fixed inputs without touching live systems or data.5/5
Policy Engines BundleMaps three policy engines over test data without model calls or live campaign execution.5/5
Audio Level RMS PortComputes the audio loudness math on test arrays without opening a microphone or capturing input.3/5

Job Computes the audio loudness math on test arrays without opening a microphone or capturing input.

Scope limit projection mechanics only, not domain-level conclusions

EvidenceComputed projectionevidence 3/5Source-faithful refactor

source intakeprovenancedrift-control

Structural Theses Finance BundleRuns a copied finance-thesis model through dated test cases with no live market data or advice.5/5

Job Runs a copied finance-thesis model through dated test cases with no live market data or advice.

Scope limit It validates only the imported source body over synthetic thesis rows. It does not claim source authority, whole-system equivalence, financial or investment-related actions, live market data, portfolio action, external model access, launch, public sharing, launch, or public sharing.

Engine Room DemoRuns proof, runtime, security, and routing demos through bounded public examples with stated limits.5/5
Backend & Governance Engines BundleExercises thirteen copied backend and governance engines over fixed public test cases.5/5

Job Exercises thirteen copied backend and governance engines over fixed public test cases.

Scope limit It validates only a public source-open bundle and bounded synthetic exercises; it is not live lineage truth, human approval authority, market/news truth, host-state truth, work log truth, external model access, source-file changes, public sharing, launch-scope decision, or whole-system equivalence.

Governance & Compiler Mechanisms BundleChecks thirteen copied governance and compiler routines against the code they were copied from.5/5
Saturation Engines BundleVerifies twelve copied engine routines and computes each failure probe from inputs, not echoes.5/5
Tool Server Pressure InventoryFlags detached helper processes and launch pressure from synthetic rows, not live hosts.5/5
Compliance Pipeline BundleConfirms six copied compliance source files carry their functions; runs one helper on sample text.3/5

Job Confirms six copied compliance source files carry their functions; runs one helper on sample text.

Scope limit validates declared public Set 8 compliance pipeline bundle contract only; no full compliance-ledger freshness, external model access, model dispatch, source-file changes, source note mutation, launch, public sharing, non-public data equivalence, or whole-system correctness

EvidenceComputed projectionevidence 3/5Source-faithful refactor

source intakeprovenancedrift-control

Live Source Drift BundleCompares four copied router and landing routines against current code to surface stale copies.5/5

Job Compares four copied router and landing routines against current code to surface stale copies.

Scope limit verified source body import only, not route authority, work log or work log mutation authority, mission-transaction execution, git staging or commit approval, source-file changes, non-public runtime export, launch, or public sharing

Release Public Wording GateFlags affirmative open-source and deployment-posture wording while allowing safe boundary notes.5/5

Job Flags affirmative open-source and deployment-posture wording while allowing safe boundary notes.

Scope limit This is lexical fixture evidence only; it is not launch-scope decision, not publishing-scope decision, not semantic NLP truth, not secret-scan coverage, and not whole-system correctness.

Generated Projection Drift RuntimeRuns each owner's own drift check on generated files and flags repairs, without changing the file.4/5Runs real tools

Job Runs each owner's own drift check on generated files and flags repairs, without changing the file.

Scope limit The bundle demonstrates owner-routed generated projection drift detection over declared artifacts, source authorities, clean-result record fingerprints, and no-write check return codes. It is not semantic drift proof, not full source registry validation, not repair authority, and not launch-scope decision.

Work & continuity (5)

Mission Transaction Work SpineRuns the real work-ledger engine on a sanitised snapshot to re-derive each change's verdict.4/5Runs real tools

Job Runs the real work-ledger engine on a sanitised snapshot to re-derive each change's verdict.

Scope limit It validates work-landing, claim, checkpoint-lane, and dependency metadata projections over fixed fixtures only; it does not mutate live ledgers or git, certify real completion, authorize broad staging without operator intent, or prove any change is actually correct or complete.

EvidenceBounded runtime computationevidence 4/5Real runtime result

agent-concurrencyworkflow-engineeringcontinuity

Durable Agent Work Landing ReplayAudits recorded work-claims so each cites files, validates before commit, and proves HEAD moved.5/5

Job Audits recorded work-claims so each cites files, validates before commit, and proves HEAD moved.

Scope limit It validates only the declared public work-landing contract over recorded rows. It is evidence for fixture-local completion mechanics, not for live Git side effects, unrelated-path staging, non-public body export, service operation, or distribution clearance.

Bridge-Continuity Sign-off ReplayReplays a paused job to prove the rules for safely resuming it hold and reject duplicate resumes.5/5

Job Replays a paused job to prove the rules for safely resuming it hold and reject duplicate resumes.

Scope limit It validates only the declared public continuity contract over synthetic fixtures; it does not run live bridge transport, use external model services, read operator HUD/browser/phase-runtime or private-memory state, prove provider or UI uptime, land work, change source files, or include launch operations.

Concurrency Mission ControlRuns copied claim-coordination code so duplicate, stale, and conflicting claims get blocked.5/5

Job Runs copied claim-coordination code so duplicate, stale, and conflicting claims get blocked.

Scope limit verified concurrency mission-control source body import only, not a live scheduler, external model access, hosted orchestration, production concurrency-safety proof, source authority, whole-system equivalence, public sharing, or launch-scope decision

EvidenceVerified source importevidence 5/5Copied source body

agent-concurrencyworkflow-engineeringcontinuity

Semantic Singleflight Dedup RuntimeReuses one result for duplicate command runs, keyed by repo state rather than the words alone.4/5Runs real tools

Job Reuses one result for duplicate command runs, keyed by repo state rather than the words alone.

Scope limit It keys and dedups command runs by a repo-state fingerprint over bounded public fixture commands only; it does not guarantee global mutual exclusion, does not replace a lock service, cannot prove cross-host correctness, and is not a job scheduler, a daemon, or launch-scope decision.

EvidenceBounded runtime computationevidence 4/5Real runtime result

agent-concurrencyoperational-disciplinededuplication