Since we published SuperQode 2.5, the work on the terminal has been about making the harness usable as a daily coding surface: connect a provider or an agent, stay in the TUI, and see what the run is doing. While that was landing, the Jev engineering guide for coding agents was published. We read it as a description of how a coding harness should handle context, evidence, tools, and execution, not as a single feature to copy. Other coding-agent harnesses have been taking similar ideas one at a time. One product adds model routing. Another adds tool routing. A third adds a narrower context trim. We decided SuperQode should line up with the guide as a whole. This work does not implement every point in that document. It implements the points we could put into the coding workflow with a typed decision, a retained artifact, and a view the developer can inspect. Automatic model, provider, and reasoning-effort switching stay outside this scope. The developer still chooses the coding model.
We have published an open-source repository so you can walk the context integration yourself: SuperQode Jev Engineering. A short TUI walkthrough sits with it. The repository has the configuration and prompts for the same workflow.
Why Jev sits beside the coding model
TypeSafe’s Jev returns structured decisions that application code can consume. The coding model can keep writing code and explanations, while Jev answers a narrower question at a defined point in the workflow. Diogo Almeida introduces that split in Introducing System One Models & Jev: supply state, ask typed questions, and use the returned decisions inside software. The official API provides three primitives. Choice picks among predefined options. Score places a state on a defined rubric. Noul returns the probability that a statement is true. The surrounding application gets a predictable interface for branching and for uncertainty.
For SuperQode, the repeated coding question was concrete. Does the next step still need the full text of this earlier tool result? A repository read, a search result, or a test log can be essential when it arrives. Several steps later, a bounded excerpt and a retrieval reference can be enough, as long as the evidence is still there to open.
From the notes to the implementation
The Jev Engineering for Coding Agents note explores a broader harness design: explicit state, context visibility, tool disclosure, retrieval reuse, and execution policy. Its cover identifies it as an independent synthesis based on design notes attributed to Diogo Almeida, not an official TypeSafe paper. A text mirror is here. We used those ideas for a focused implementation on the existing Core harness, PiPy sessions, WorkOrders, permissions, and evaluation. The priorities were better context, reusable evidence, and reliable execution.
Keeping evidence and putting that evidence into every model request are separate jobs. The hosted context integration stores permitted, sanitized tool output in a durable SQLite artifact store. Each artifact has a stable reference, a content digest, an ownership scope, and provenance. When enforcement selects an excerpt, the model-facing context gets a bounded preview and a retrieval pointer. The retained artifact stays available through read_context_chunk. Retrieval returns bounded pages, checks current permissions again, and returns an explicit error when evidence is missing or denied. It does not rerun the originating tool. Original evidence here means the retained, permitted representation. Sanitization may remove sensitive material before storage. The PiPy session archive stays intact, because context selection changes the projection prepared for a model call rather than the stored history. Details are in the 2.5.5 context notes.
A small decision, with local protections around it
The selector needs more than a tool name and a byte count. The shared context policy sends bounded state: the task, recent context, and previews of eligible evidence, including matching or error spans. The current integration uses Noul to estimate whether the full evidence is still needed. SuperQode then keeps the full eligible tool result, or replaces it with a bounded excerpt and a retrieval reference. The shared policy excerpts only when the returned probability that full evidence is needed is below 0.2. Any other valid probability keeps the full evidence. An invalid answer, a timeout, an unavailable client, or an exhausted call allowance also keeps it. A typed answer makes the branch predictable. It does not prove that every relevance judgment is right. Thresholds still need evaluation on representative coding tasks. TypeSafe’s own confidence guidance says the same about domain-specific testing.
Jev only selects among candidates SuperQode has already bounded. The policy starts with older, larger text tool results. Recent evidence, error results, and ineligible message structures stay off that path. User requests and instructions are not replaced. The host still authorizes every read, retrieval, and command. A relevance decision grants no new permission. Because the selector is hosted, the bounded decision state is sent to the configured TypeSafe service. Redaction reduces exposure, and developers still decide which task and evidence that service is allowed to receive.
Decisions are persisted and reused while their inputs stay valid. A small plain assistant continuation can keep the previous selection. A change to the task, the instructions, the candidate evidence, the relevant history, or context pressure invalidates it. Shadow mode records the proposal and leaves baseline tool output in place. Enforce mode applies eligible excerpt decisions. Off disables the shared policy. A smaller character count in shadow mode is an observation about a possible projection. It is not a measured token saving, and a smaller prompt can still produce a worse answer. Regression coverage includes continuation, restart, invalidation, uncertainty, and selector failure. Those tests check the implementation. They do not yet measure decision quality on live coding tasks. See test_context_policy.py.
Inspect it in the TUI
The same terminal is where you read the decision. Enter :context evidence. The inspector shows the mode, the selector, whether a cached decision was reused, actual and proposed character counts, selector calls, and each evidence row. Open a row for the action, reason, probability, and proposed excerpt. Open original retrieves the retained artifact, and Next page continues through it. A recorded call attempt is not the same as a usable decision. A fallback such as scorer_failed means the selector did not return one, and the evidence stays retained. The inspector covers context preparation SuperQodecontrols. External harnesses keep their own visibility limits. The widget is context_evidence.py.
Try the example
The public example targets SuperQode 2.5.5 and uses a synthetic runbook. The usual install remains curl -fsSL https://superqode.dev/install.sh | sh. The pinned example below uses uv so the walkthrough matches 2.5.5.
uv tool install superqode==2.5.5
git clone https://github.com/SuperagenticAI/superqode-jev-engineering.git
cd superqode-jev-engineering
export TYPESAFE_API_KEY="YOUR_OFFICIAL_TYPESAFE_KEY"
sq
Inside the TUI, use :connect to choose the coding provider and model. Then enter these one at a time:
:systemone shadow
:harness use ./demo-shadow.yaml
:pipy new
Send:
Use the read tool to read the entire runbook.txt file.
Then reply only "Runbook read". Do not modify files.
After the response, send:
Explain stable evidence references in one sentence.
You do not need the detailed historical runbook. Use no tools.
The earlier read is now eligible for context selection. Open :context evidence, inspect the decision, and choose Open original. The final page of the retained runbook contains END_MARKER=context-original-intact. To try enforcement, enter:
:harness use ./demo-enforce.yaml
:pipy new
Repeat the two prompts and inspect the result. If Jev selects an excerpt, enforcement reduces the actual context characters and retrieval still works. If it selects keep, the full evidence stays. Both are valid. The example bounds candidates and selector calls per run. It is not a total session spending cap. Jev calls and coding-model calls bill on their own accounts. The repository guide has configuration and troubleshooting.
Watch Demo
WorkOrders, recovery, Pi Durable, and Monty
Context retrieval helps one conversation. WorkOrder evidence reuse helps a later worker. An investigator can publish findings as addressable artifacts. An implementer or reviewer can inspect assigned predecessor evidence instead of starting the investigation again. SuperQode records source hashes, repository identity, relevant environment fingerprints, and predecessor lineage. Freshness is current, stale, or unknown. Verification is separate. An agent report starts as reported. Matching source hashes do not turn its conclusions into verified facts. A tool receipt is an observed output, and its meaning can still need a check. Acceptance checks and candidate review still run. See the WorkOrder evidence notes.
Opt-in PiPy recovery inside supported WorkOrders records complete model responses and tool outcomes before advancing. Restart can reuse a committed outcome when invocation identity, workspace, configuration, and current policy still permit it. An interrupted operation with an unknown outcome needs reconciliation. Uncertain unsafe operations and model requests are not repeated automatically. :work view ID shows dependencies, acceptance, evidence freshness and verification, recovery status, the candidate diff, acceptance output, and human approval. Ordinary interactive session resume does not create these checkpoints. Workspace verification can detect a change. It does not recreate a deleted file, restore an environment, or reattach a shell. See PiPy recovery.
Pi Durable remains a reference for durable agent execution. SuperQode WorkOrders are the delivery contract around that execution: dependencies, evidence, acceptance, recovery review, and a human decision on a concrete candidate. PiPy recovery reuses committed model and tool outcomes. It is still narrower than full Pi Durable compatibility, and matched live comparisons are still open. Monty is the other bounded path: Python programs that compose admitted tools and process results. Inside a supported WorkOrder, SuperQode can checkpoint the suspended interpreter and inject committed outcomes after restart. The PiPy Monty integration does not expose general filesystem access, shell execution, writes, or arbitrary third-party imports inside the program. Its checkpoints cover the program. They do not cover every model request or shell process. Hosted PiPy runs with process permissions and contextual call and result policy. It does not provide an interactive tool-approval stack or an OS sandbox. The boundaries are in tool composition and recovery.
What is in place, and what we still need to measure
SuperQode now has an opt-in path from a typed context decision to retrievable evidence and a visible run. You can inspect a live decision, compare the proposed and enforced representation, open omitted evidence, and see why a fallback happened. WorkOrders add lineage, freshness, acceptance, and conservative recovery for dependent tasks. The next measurement holds the coding model fixed and compares the same tasks with selection off, in shadow, and enforced. That comparison should record correctness, provider-reported tokens, completion time, retrieval, recovery, and total reported spend, including selector overhead. The TUI example shows the mechanics. It does not establish better coding quality, lower total cost, or a win over Pi. If you want to try the workflow or send a reproducible failure, start with SuperQode Jev Engineering on GitHub.
- Product: superqode.dev
- Docs: docs.superqode.dev
- PyPI: superqode
