Area · 20 components
Agent reliability & safety
Source-open replays of agent failure modes as inspectable specimens.
Components
Each card below is one component. It states what the component does, the evidence behind that claim, and its scope limit: the line where the claim stops and nothing further is proven.
Agent Completion Faithfulness AuditRuns real git and pytest on a sample repo so wrap-up claims state only what the evidence proves.4/5Runs real tools
Bounded Autonomy Campaign PacketDrafts proposed work from coverage gaps and proves it cannot repair or rewrite the code itself.4/5Runs real tools
Provider Context Recipe Budget PolicyRuns the real context harness to measure assembled byte sizes and check each bundle fits its budget.4/5Runs real tools
Job Runs the real context harness to measure assembled byte sizes and check each bundle fits its budget.
Scope limit It validates context-budget projection mechanics (byte ceilings, ordered section fill, omitted-section manifests, deliverable routing, and digest-checked source-body imports) only. It excludes provider/API calls, run Lean/Lake, expose or carry proof or oracle truth-side material, assert theorem or domain-level conclusions, or include launch operations.
Cold Evaluation Honesty BundleRuns a copied route-quality simulator and checks its all-B scorecard against the original code.5/5
Job Runs a copied route-quality simulator and checks its all-B scorecard against the original code.
Scope limit verified cold-eval source body import only, not a live benchmark, navigation truth, source authority, external model access, whole-system equivalence, public sharing, or launch-scope decision
Validator Checker BundleRuns the real validator code over public examples so its safety checks stay inspectable.5/5
Job Runs the real validator code over public examples so its safety checks stay inspectable.
Scope limit It validates only the imported validators.py source body and its checker membrane. It does not claim source authority, a full validator-suite proof, whole-system equivalence, launch, hosted-public status, public sharing, external model access, or source-file changes.
Secondary Runtime Source BundleRuns eight trace, graph, and market engines on test rows without fetching live markets.5/5
Job Runs eight trace, graph, and market engines on test rows without fetching live markets.
Scope limit verified source body import only; no browser/session export, wallet authority, live market data, investment-related actions, external model access, source-file changes, whole-system equivalence, public sharing, launch, semantic-truth, or whole-system correctness claim
Agent Benchmark Integrity Anti Gaming ReplayValidates a synthetic benchmark-integrity record and flags the contamination cases it declares.3/5
Job Validates a synthetic benchmark-integrity record and flags the contamination cases it declares.
Scope limit It authorizes only bounded public runtime validation over copied source-open pattern provenance bodies and metadata-only benchmark-integrity replay rows; it does not establish any benchmark or SWE-bench score, agent capability, external model service, live-repo mutation, private/oracle/hidden-gold body access, product progress, or launch-scope decision.
Monitor Evidence-Boundary ReplayReplays honest and deceptive agent runs and flags any verdict missing its declared backing evidence.3/5
Job Replays honest and deceptive agent runs and flags any verdict missing its declared backing evidence.
Scope limit Bounded public runtime validation over copied source pattern bodies, sanitized dogfood trace slices, recomputed monitor-verdict spans, source-artifact evidence refs, digest/metadata-only/non-public-state gates, and negative cases only; no live agent execution, monitor product performance, control-eval score, safety-validation, benchmark, provider-call, source-file changes, launch, public sharing, or product authority.
Sabotage-Monitor Contract ReplayAudits a hidden-goal catch claim for the steps, suspicion scores, and counterfactual it needs.3/5
Job Audits a hidden-goal catch claim for the steps, suspicion scores, and counterfactual it needs.
Scope limit Bounded public runtime validation over copied source pattern bodies, sanitized dogfood trace slices, recomputed sabotage/scheming monitor spans, source-artifact evidence refs, digest/metadata-only/non-public-state gates, and negative cases only; no live sabotage, live agent execution, exploit instruction, account secret/account, private-reasoning, harmful-payload, monitor-product-performance, deployment-risk, benchmark, provider-call, source-file changes, launch, public sharing, or product authority.
Agent Memory Temporal Conflict ReplayReplays a memory edit-and-delete to show stale facts get flagged before they sway an answer.3/5
Job Replays a memory edit-and-delete to show stale facts get flagged before they sway an answer.
Scope limit It validates the projection mechanics of a synthetic memory fixture only — that the required refs, decisions, paired replays, negative cases, and secret-exclusion scan line up and that result records are metadata-only. It does not claim live-memory product quality, judge whether memory decisions were domain-correct, treat memory recall as source authority, adopt active injection, export private transcripts, use external model services, change source files, or include launch operations.
Memory-Poisoning Quarantine Policy ReplayReplays a recorded memory-tamper case, checking its declared quarantine, block, and delete steps line up.3/5
Job Replays a recorded memory-tamper case, checking its declared quarantine, block, and delete steps line up.
Scope limit It only checks the structural shape and internal consistency of a synthetic memory-security policy projection recorded as JSON. It does not run or validate any real memory store, does not itself quarantine, delete, or re-run anything, and does not establish that any system actually resists poisoning. It exports no private memory bodies or transcripts, calls no providers, mutates no source, produces no benchmark claims, and excludes launch (all scope limit flags are hardcoded false).
MCP Tool-Authority Policy ReplayAudits a recorded tool-use log to confirm each action was scoped, approved, undoable, and fenced.3/5
Job Audits a recorded tool-use log to confirm each action was scoped, approved, undoable, and fenced.
Scope limit It only checks that the tool-authority evidence in a recorded bundle (scopes, approvals, rollbacks, instruction/data splits, cold replays, redaction, and the expected abuse-case failures) is present and internally consistent. It does not run tools or authorize live MCP/account access, account secret or payload export, treating tool output as instruction, source-file changes, benchmark safety scores, or launch, and it makes no claim that the underlying tool-use policy is domain-correct.
Belief-State Reward Bundle ReplayChecks that each step reward in a recorded run cites a declared verifier-feedback row, not a trick.3/5
Job Checks that each step reward in a recorded run cites a declared verifier-feedback row, not a trick.
Scope limit It only checks that the projection's accounting lines up under its own schema rules over recorded synthetic fixtures; it excludes hidden-reasoning export, RL training, hidden gold or neural-judge-only labels, benchmark-performance claims, external model access, source-file changes, or launch, and proves nothing about real-world reward, live agent behavior, or domain-level conclusions.
Sandbox-Policy ReplayMaps sandboxed agent actions to show each was approved or blocked before running, then rolled back.3/5
Job Maps sandboxed agent actions to show each was approved or blocked before running, then rolled back.
Scope limit It validates the projection / trace-refactor mechanics over a synthetic fixture only; it excludes live sandbox escape, secret or account secret handling, live network access, host filesystem mutation, executable payload export, raw environment export, external model access, security benchmark claims, source-file changes, or launch. A pass proves the projection boundary and trace-refactor mechanics for this contract, not real sandbox security, exploit resistance, or whole-system safety.
Prompt-Injection Flow-Policy ReplayReplays an agent run to show untrusted text was gated before any sensitive action, leaking no secret.3/5
Job Replays an agent run to show untrusted text was gated before any sensitive action, leaking no secret.
Scope limit Passing result records only show this projection satisfies the named information-flow contract over synthetic, redacted, metadata-only rows; they do not prove general prompt-injection robustness, benchmark performance, live account/tool/provider safety, hidden-message handling in a real system, source-file changes, or launch-scope decision.
Vulnerability Patch-Proof ReplayChecks a fixed-bug evidence chain and re-runs three small real security checks; no real attack material.3/5
Job Checks a fixed-bug evidence chain and re-runs three small real security checks; no real attack material.
Scope limit It validates only the projection/evidence-chain mechanics of a synthetic replay: structural presence, cross-reference consistency, declared boolean flags, and the secret/live-access exclusion scan. It executes small regression witnesses but performs no real vulnerability discovery and makes no judgment of real-world security or fix correctness. It excludes live-target testing, real CVE exploitation, weaponized payloads, account secret handling, network exfiltration, actionable exploit steps, external model access, source-file changes, benchmark security scores, launch, or any whole-system security claim.
Agent Route Observability RuntimeRecomputes an agent run's route-compliance score and anti-pattern flags with real trace-analytics code.5/5
Job Recomputes an agent run's route-compliance score and anti-pattern flags with real trace-analytics code.
Scope limit It validates only public, recorded trace-feedback metadata and regression fixtures; it does not inspect live operator state, certify or prove runtime behavior, read model-output data, mutate the work log, authorize pattern assimilation, or include launch operations.
Bridge Campaign DAG ValidationRuns shape checks on a fan-out work plan for unique steps, dependencies, and no cycles, not the plan itself.4/5Runs real tools
Job Runs shape checks on a fan-out work plan for unique steps, dependencies, and no cycles, not the plan itself.
Scope limit The bundle validates a public bridge-campaign DAG contract and provider worker ceiling. It is not a dispatcher, not a live multi-agent run, not a provider safety proof, and not launch-scope decision.
Metabolism Queue ReconciliationRuns a scratch-database model of a durable job queue and flags impossible states for a human to review.4/5Runs real tools
Job Runs a scratch-database model of a durable job queue and flags impossible states for a human to review.
Scope limit The bundle demonstrates a synthetic SQLite durable queue, lease recovery, blackboard claim-event projection, and cold-start reconciliation taxonomy. It is not a live non-public runtime export, not an agent dispatcher, not external model service, not ambiguous auto-repair, not a distributed database, and not launch-scope decision.
Egress Self-Compliance AuditRuns phrase-membership checks on an agent's own replies for three self-policing slips, not deep meaning.4/5Runs real tools
Job Runs phrase-membership checks on an agent's own replies for three self-policing slips, not deep meaning.
Scope limit The bundle evaluates egress text through explicit phrase-membership policy. It is not taint analysis, prompt-injection defense, sandboxing, information-flow control, or launch-scope decision.
Source refs
Built from public source refs, with each input path recorded for provenance.