RLM + A2A in SuperQode: Building a More Durable Recursive Coding Harness

We recently released SuperQode with a native Recursive Language Model harness. Getting that harness working was an important step. Making it useful for longer coding tasks required more work: reliable communication between agents, recovery after interruptions, clearer execution boundaries, and visibility into the cost of recursive inference.

An RLM can break a problem into smaller questions and delegate work to other model calls or agents. That raises practical questions. Where do their results live? Can a child agent talk to its parent or a sibling? What happens if the terminal disconnects? Can the root accidentally spend its allowance several times by spawning children with independent budgets? And when an external agent says it has finished, who verifies the result?

Those questions shaped our recent releases. In a recent version we added optional Agent2Agentprotocol delegation inside native RLM, retained local agent inboxes, Monty resource limits and persistent checkpoints, and stronger recovery around remote tasks. We also added durable history, a shared inference allowance, and experimental Python+Bash and selective-observation profiles. The model decides how to investigate. The harness owns admission, persistence, permissions, accounting, and verification.

Research that shaped the work

Our starting point was research into how other harnesses handle this space. Prime Intellect’s Prime Agent was especially useful. Its persistent Python environment lets the model compose operations programmatically, work with retained agents, and communicate across sessions. That makes the execution environment part of the reasoning process.

There is a useful distinction. In the Prime revision we inspected, agent communication used custom session messaging through its host and supervisor. Our audit did not find an integration of the open A2A protocol on that path. We took inspiration from the retained-agent design and connected SuperQode’s existing A2A implementation to native RLM. The source revision and findings are in our A2A research notes. The shipped APIs and recovery commands are in the RLM A2A guide.

Alex Zhang’s writing on harnesses as compositional generalizers helped frame the engineering question: a harness can shape a difficult task into smaller computations the underlying model can handle. We kept the RLM’s ability to inspect and decompose context, then hardened the work that follows.

The original Recursive Language Models formulation treats long input as part of an external environment the model can examine programmatically and recursively query over selected portions. In SuperQode, repository context is data in a persistent Python environment. The model can search it, select files, split material into chunks, ask focused semantic questions, and combine answers. On supported profiles it can start local coding children. Large intermediate results stay available through handles and bounded reads, so later model calls need not receive the entire result again.

We kept three distinct operations. llm_query(...) asks a focused semantic question over selected context. rlm.run(...) starts a local recursive child coding session when the profile allows it. a2a.start(...) delegates bounded work to an explicitly configured independent agent. A small question can stay a direct model call. A repository investigation can become a local child. A specialist review can go to an external agent when that route is enabled. Direct inference stays separate from remote agent tasks, including their network and lifecycle overhead.

A2A as an opt-in capability of the RLM program

A2A supplies the external delegation contract. An Agent Card describes a peer’s interfaces, capabilities, and authentication requirements. A request can return an immediate message or a task with its own lifecycle. The caller can inspect that work and exchange further input without access to the remote agent’s internal tools or prompt. The model invokes that capability from its RLM program. The host executes the protocol and returns a local handle. No model-weight change or provider inference API change is required.

After enabling and configuring a reviewer peer, the model can select a small evidence bundle and submit it:

PYTHON
selected = context.select('src/auth*.py')
bundle = [
    {'path': c.path, 'start': c.start, 'end': c.end, 'text': c.text}
    for c in selected.chunk(size=12000)
]

review = a2a.start(
    peer='reviewer',
    task='Audit these authentication paths',
    context=bundle,
    request_id='auth-audit-v1',
    required=True,
)

review.wait(timeout=20)
review.status()
findings = review.read(size=4000)

The handle supports progress inspection, bounded result retrieval, replies to clarification requests, cancellation requests, and follow-up work. Waiting for 20 seconds does not mean the remote task has stopped. A peer that refuses cancellation is not a successfully cancelled task. File URLs from a peer remain references; SuperQode does not fetch them implicitly.

Much of this work was making those distinctions survive beyond a happy-path demo. We hardened the A2A client to distinguish direct messages from tasks, preserve authentication-required and input-required states, and retain artifact information. The delegation manager records admission before sending, stores bounded results, and tracks work across terminal detachment and worker restart.

A stable request ID returns the existing local handle when the inputs match. Reusing it with changed inputs is refused. If a remote task ID is known, recovery can ask the peer for its authoritative state. If a send may have reached the peer but its acknowledgment was lost, the outcome stays unknown. SuperQode does not automatically submit the work again. A2A message IDs alone cannot establish exactly-once execution.

Local recursive agents use a separate path. Parents, children, and siblings within the same recursive root can exchange retained inbox messages. Reads do not consume a message; the receiver acknowledges after handling it. Delivery IDs support idempotent delivery. Unrelated roots cannot read those inboxes. Delivering a message does not wake a model or create an inference charge. Explicit child continuations admit new work against the retained conversation while preserving the original run’s terminal record.

Verification stays with SuperQode. A remote agent’s completed status is evidence. Its explanation or patch still needs local inspection and the configured completion gates. Required remote work and active local children must be resolved before the root completes. WorkOrders keep their acceptance path, with remote receipts treated as unverified evidence.

Host, Docker, and Monty

We support three execution boundaries because they serve different workloads. Host is for local coding with installed tools and dependencies; generated Python has host-process permissions, and harness controls are guardrails. Docker is for coding and command execution inside a container; the image must provide the tools and dependencies, and its configuration determines access. Monty is for restricted Python context processing, history reads, and semantic analysis. Our Monty profile has no shell, repository writes, or coding children.

Monty is Pydantic’s Rust-based Python interpreter. Host functions expose specific capabilities to generated Python while keeping their implementation outside the interpreter. In our Monty profile, allowed model and A2A operations run through explicit host bridges. The interpreter receives opaque handles and bounded data. Credential resolution and remote transport stay on the SuperQode host. The Docker bridge follows the same idea: enabling a specific A2A capability can allow host-mediated delegation without granting general container network access.

Monty also gained resource limits and persistent checkpoints, with compatibility and integrity checks on restore. Successful feeds can retain state for a replacement worker. Its memory accounting is not a process RSS ceiling, and remote or inference waits have separate controls.

Cost, hosted peers, and the shared ledger

Cost shaped routing from the start. A2A is off by default. Installing a key does not enable a route or paid fallback. Users explicitly configure peers, credential environment references, task limits, concurrency, and payload and result allowances. Credentials are referenced by name rather than written into the profile.

User-owned peers and SuperQode-hosted specialists are separate choices. Speaking the open protocol does not require SuperQode-hosted inference. Hosted execution needs explicit enablement, eligible customer authorization, and a credit allowance. We added durable customer credit reservations and settlement for bounded specialist tasks; reconnecting to existing work does not charge again. That implementation supports a controlled single-host pilot. Paid execution remains disabled on the public Cloud Run catalogue. Multi-instance hosted execution needs a shared durable backend. Native RLM features can be used without a Cloud Run change or an automatic hosted route.

We also extended accounting to the native recursive family. Root inference, coding children, and semantic queries share a durable inference ledger and call allowance. Starting another child or restarting a worker does not create a fresh allowance. Call admission is atomic. Token and USD thresholds use provider-reported usage, so a current request can cross the threshold before later calls stop. With either threshold enabled, concurrent inference waits for a live request to settle. Missing usage or prices remain unknown and block further admission under the corresponding threshold. Interrupted inference needs explicit reconciliation. Remote A2A usage and hosted credits stay separate from this native USD accounting.

This makes spending easier to inspect and constrain. It does not establish that recursion, A2A, or Monty lowers total task cost. Monty may reduce execution overhead for suitable programs; inference cost still depends on how many calls the model makes and how much context those calls receive.

Durable history and experimental profiles

Longer tasks need retrievable history. SuperQode now adds branch-scoped history that includes original tool output and messages from before compaction. The model can search that history and read a bounded page by handle:

PYTHON
matches = history.search('failing assertion', limit=10)
page = history.read(matches[0]['id'], start=0, size=4000)

An experimental selective-observation profile uses compact receipts, previews, and history references in the provider-facing conversation. Original session records remain intact. User instructions, roles, and tool-call and result identities are preserved. Showing less output can help focus a call. It can also hide information the model needs. That tradeoff is still to evaluate.

We also added an experimental Python+Bash surface. Python remains useful for selecting evidence, managing handles, and recursive computation. Bash makes builds, tests, and existing command-line tools easier to use. Both operate in the selected host or Docker environment and share the workspace and command-job ledger. Host execution needs POSIX process groups; Windows users should choose Docker for this profile. Native Python remains the default.

Bash jobs have durable receipts, bounded output, stable request IDs, deadlines, and process-group cancellation. Managed mutating commands serialize across the recursive family. Interrupted jobs retain an unknown outcome and a workspace lease until their effects are verified. Explicit read-only declarations support concurrency. They do not enforce filesystem isolation; raw host Python and legacy shell operations can bypass this coordination.

Using it in the terminal

Open :connect, choose Connect a harness with your model, then RLM. Select host, Docker, Monty, or the experimental profiles, and configure optional A2A separately. :rlm settings exposes the tool surface, boundary, observation strategy, and inference limits. Operational commands let you inspect the active profile, budget, history, command jobs, remote delegations, and usage while the resident worker is running. Saving a project profile does not silently change that worker’s policy. Starting with different settings creates a new session. The profile guide documents the choices and commands.

Try the RLM + A2A demo in the TUI

You can reproduce the pattern above in the SuperQode TUI with the controlled examples/rlm-demofixture. The unfinished function is build_release_health in deploy_audit/report.py. The runbook, parser, deployment fixture, and unit tests define the correct behavior: an older success must not hide a newer failure, duplicate events must not create duplicate services, and a successful deployment taking more than 300 seconds must be marked slow. This is an example project for inspection and verification, not a claim that the agent is fixing a live production incident.

You will drive a free NVIDIA Nemotron model on OpenRouter through SuperQode’s native RLM. A local recursive child reviews the parser contract. An independent A2A reviewer runs as a read-only Monty harness. Docker runs the tests. The reviewer needs its own terminal, because the TUI connects to an A2A server and does not launch that server for you.

1. Prepare

Install the recent version of SuperQode with the reviewer dependencies:

SHELL
uv tool install --force --python 3.12 'superqode[a2a,monty]'

Start Docker, confirm it is running, and pull the Python image:

SHELL
docker info
docker pull python:3.12-slim

From a checkout of the SuperQode repository, move into the demo directory:

SHELL
cd examples/rlm-demo

If you have run the demo before, restore the unfinished function:

SHELL
git show HEAD:examples/rlm-demo/deploy_audit/report.py > deploy_audit/report.py

2. Terminal A: start the A2A reviewer

Leave this running:

SHELL
export OPENROUTER_API_KEY='YOUR_OPENROUTER_KEY'
superqode serve a2a \
  --spec rlm-monty.yaml \
  --provider openrouter \
  --model 'nvidia/nemotron-3-ultra-550b-a55b:free' \
  --host 127.0.0.1 \
  --port 8000 \
  --working-dir .

3. Terminal B: check the Agent Card and open the TUI

In a second terminal, in the same demo directory:

SHELL
export OPENROUTER_API_KEY='YOUR_OPENROUTER_KEY'
curl -fsSL http://127.0.0.1:8000/.well-known/agent-card.json
superqode --harness core

The curl command should return the reviewer’s Agent Card.

4. Choose RLM and the model

In the TUI, open the connect menu:

TERMINAL
:connect

Choose Connect a harness with your model, then RLM, then Configure RLM profile and budgets. Use these settings, then choose Start new session:

  • Tool profile: Python + Bash (experimental)
  • Execution environment: Docker
  • Model observations: Conversation transcript
  • Shared model-call allowance: 20
  • Token threshold: 0
  • USD threshold: 0

For the model, choose BYOK, then OpenRouter, then NVIDIA: Nemotron 3 Ultra (free), ID nvidia/nemotron-3-ultra-550b-a55b:free. If the model menu does not open, run this and press R to refresh:

TERMINAL
:connect models

5. Enable A2A

Open the A2A settings:

TERMINAL
:rlm a2a

Use these settings:

  • Enable A2A routing: on
  • Paid hosted peers: off
  • Hosted credit limit: 0
  • Max remote tasks: 2
  • Simultaneous remote tasks: 1
  • Peer alias: reviewer
  • Agent URL: http://127.0.0.1:8000
  • Description: Read-only review of release-health rules and tests
  • Key environment variable: empty
  • Agent skill: superqode-harness
  • Paid hosted service: off
  • Tariff: 0

Choose Add / update peer, confirm reviewer is listed, then choose Start new session. Check the session:

TERMINAL
:rlm profile
:rlm routing
:rlm peers

Confirm Python+Bash, Docker, A2A on, and the reviewer peer.

6. Paste the task

Paste this task into the main conversation:

TASK
Implement build_release_health in deploy_audit/report.py.

Read RUNBOOK.md, INCIDENT.md, deploy_audit/parser.py,
fixtures/deployments.jsonl, and tests/test_report.py.

Use one local rlm.run child to review the parser contract.
Use one required a2a.start delegation to peer="reviewer" for
an independent read-only review of the rules and tests.
Send bounded relevant context.

Wait for both reviews and read their results before implementing.
Leave the tests unchanged.

Run python3 -m unittest discover -s tests -v using commands.run.
Wait for completion and read the output.

Report the implementation, local child ID, A2A delegation ID,
test command job ID, and actual test result. Keep responses concise.

7. Show the results

TERMINAL
:rlm agents
:rlm delegations
:rlm jobs
:rlm job read ACTUAL_JOB_ID stderr 0 4000

The delegations list should show a completed delegation with peer=reviewer. Replace ACTUAL_JOB_IDwith the test job’s ID from the jobs list. Python’s unittest report lands on stderr. The intended successful output is Ran 3 tests and OK.

Watch Demo

For the full command surface, recovery notes, and profile details, use the RLM A2A guide, the profile guide, and the example README.

Continual harness, and what we measured

We also read Seth Karten’s Continual Harness, which adapts prompts, skills, subagents, and memory from past trajectories during a continuing run. That raises a clear question about agents improving their own harness. We have not shipped autonomous harness refinement in these releases. Durable history and explicit verification are useful foundations for exploring it later. Letting an agent change its operating rules raises correctness questions around permissions, credentials, spending policy, and acceptance authority. Those boundaries need to be clear first.

We built a coding-profile pilot runner to compare experimental options. It uses fresh workspaces, fixed model selection, independent graders, and per-attempt usage records. Docker attempts are graded inside Docker. The bundled pack has 12 small single-file tasks. Offline mode uses scripted edits and checks execution plumbing across the profile presets. It does not measure model quality or demonstrate savings. Representative live coding tasks and repeated measurements remain necessary. The pilot example documents that distinction.

What these releases put in place

Together, these releases make SuperQode’s RLM more inspectable and more deliberate about recursive work. The model can program its investigation, retain evidence, communicate with local agents, and optionally delegate to independent A2A peers. The harness preserves task identity, execution policy, budget state, and the verification decision around that work.

Install:

SHELL
curl -fsSL https://superqode.dev/install.sh | sh
sq

Already installed: superqode update, or uv tool upgrade superqode. The recent version of SuperQode is on GitHub and PyPI. Product: superqode.dev. Docs: docs.superqode.dev.

P.S: Originally posted at Superagentic AI Blog here