Skip to main content

Command Palette

Search for a command to run...

Jev Engineering for Coding-Agent Harnesses

Updated
•34 min read•View as Markdown
Jev as a decision layer: harness state flows into Jev, which routes to traditional ML, an LLM, tools and policy, or a human reviewer

TL;DR

  • Jev is TypeSafe AI's "System One" model — not a chat model, not a replacement for the LLM in a coding agent. It returns bounded, typed judgments (Choice, Score, Noul) that software can consume directly.

  • The useful architecture is: code owns the workflow, the coding LLM generates, and Jev supplies narrow semantic judgments at selected control points (relevance filtering, tool routing, prompt-injection detection, policy gating, escalation to a human).

  • This article synthesizes ten harness-engineering methods built around that division of labor, sourced from Diogo Almeida's published design notes and TypeSafe's documentation.

  • The goal is a coding-agent harness that is cheaper, more explainable, and more governable than "one giant LLM prompt with a loop."

Introduction to Jev

Coding agents are usually described as if the model were the whole product. Give a frontier language model a long prompt, expose a terminal and a few file tools, add a loop, and hope that the model keeps the task, constraints, evidence, and safety policy straight.

That design can work. It can also become expensive, opaque, and difficult to govern.

Jev suggests a different division of labor.

This entire article is based on the work of Diogo Almeida, founder of TypeSafe and creator of Jev. The ten harness-engineering methods below synthesize his published design notes on building coding-agent harnesses around Jev, including the "no KV cache" thought experiment, the routing arithmetic, the token-share findings, and the tiered-disclosure and command-gating patterns. Credit for that underlying thinking belongs to him; this article organizes, cross-checks, and sources it against TypeSafe's official documentation.

Jev is TypeSafe AI's first public System One model. Jev is not a chat model, and it is not a replacement for the LLM inside Claude Code, Cursor, GitHub Copilot, Codex, or another coding agent. TypeSafe's own guide is explicit: Jev does not generate text, write code, hold a conversation, or choose its own next action. It evaluates structured state against typed questions and returns bounded decisions that software can consume directly.

That makes Jev potentially useful inside the harness around a coding LLM.

The LLM can still investigate, explain, and write code. The harness can retain state, execute tools, enforce permissions, and verify outcomes. Jev can supply narrow semantic judgments where deterministic code is too brittle:

  • Which context chunks are relevant to this task?

  • Which skill or tool best matches the request?

  • Does a retrieved passage contain prompt injection?

  • Does a proposed command conflict with policy?

  • Which handler should receive the next step?

  • Is the evidence strong enough to act, or should the system ask for review?

The useful architecture is not “Jev replaces the coding agent.” It is “code owns the workflow, the coding LLM generates, and Jev supplies typed judgments at selected control points.”

Why Jev exists

TypeSafe's thesis is that chat models and machine-consumed decisions are different workloads. Chat models are optimized to generate useful strings for people. Jev is designed to return bounded values, probability distributions, and uncertainty signals that software can inspect without parsing prose.

TypeSafe calls Jev a System One model, borrowing the name from the fast, focused side of the System 1/System 2 distinction. In this programming model:

  • application state is passed explicitly as text or structured JSON

  • questions declare the judgment and its valid answer space

  • Choice selects among bounded alternatives

  • Score places state on ordered, descriptive levels

  • Noul returns the probability that a proposition is true

  • independent questions over the same state can be evaluated together

  • application code decides how probabilities become actions

This makes Jev closer to a semantic decision primitive than a conversational assistant. It can help code answer questions that ordinary if statements cannot answer from raw text, while preserving a bounded output contract.

How Jev is trained, and what TypeSafe claims about cost and speed

TypeSafe's launch article describes a training method it calls Reinforcement Learning for Calibrated Decisions (RLCD), contrasted with the RLHF and RLVR methods used to train conversational LLMs. Where RLHF optimizes for responses human raters prefer, TypeSafe states RLCD optimizes for epistemically honest, calibrated probabilities on bounded System One tasks. Jev also uses a parallel sampler: instead of generating one token at a time, it evaluates all questions against a state in a single pass.

TypeSafe's own published figures, current as of the September 2026 launch, include:

  • Latency: an end-to-end response time of roughly 70–500ms per call, versus 3–329 seconds TypeSafe attributes to frontier LLMs on comparable tasks.

  • Cost: input tokens priced at $0.042 per million tokens (about $42 per billion tokens), with output pricing described as unmetered because Jev does not generate free text.

  • Workflow evaluations: TypeSafe reports up to roughly 193.6x faster and 444.6x cheaper results on its own published "workflow eval" methodology, which compares Jev's typed decisions against an average of large reference LLMs across four workflows (security incidents, agent-trace observability, invoice processing, and customer service).

  • Type safety: TypeSafe states Jev cannot produce a schema-mismatched output, because valid outputs are defined in advance and enforced structurally, rather than validated after the fact the way free-text LLM output is.

These are TypeSafe's own self-reported claims, published alongside methodology notes and caveats on the same page (for example, that its workflow evaluation harness and prompts were authored by TypeSafe's own team, and that its comparison baseline is an average of two specific frontier models). Treat them as vendor-published evidence to verify against your own workload, not as independently audited benchmark results.

TypeSafe's documentation also gives explicit guidance on question design: keep each Choice, Score, or Noul question atomic, meaning it should ask one well-scoped judgment, similar to a determination a knowledgeable person could make in a few seconds given the right context. If a judgment would require weighing multiple independent factors, the guidance is to decompose it into separate questions and combine the results in code, rather than asking Jev to do the weighing implicitly. This directly supports the harness-side composition patterns developed later in this article.

Why it matters for coding-agent harnesses

A coding agent contains many decisions that are narrower than code generation itself: select context, rank skills, detect risky content, classify intent, score a patch, or decide whether uncertainty requires review. Those decisions can be separated from the generative model and made visible to the harness.

The result is not a fully autonomous Jev coding agent. It is a hybrid system:

Component Primary responsibility
Coding LLM Open-ended investigation, explanation, and code generation
Jev Narrow typed judgments with probabilities
Harness code State, thresholds, routing, retries, and composition
Policy engine Deterministic authorization and prohibited actions
Tools and sandbox Scoped execution and side effects
Evaluators Tests, regression checks, and outcome verification
Human reviewer Approval for consequential or uncertain actions

Ten methods for building explicit, typed, confidence-aware agent systems

This article develops ten engineering methods for that architecture. Some originate in the creator's harness framing. Others are supported directly by TypeSafe's documentation and cookbooks. Where the evidence differs, the distinction is stated explicitly.


Source boundary: what is official and what is architectural interpretation

Jev is new, and its public material is evolving quickly. A responsible design should separate confirmed product behavior from proposed harness architecture.

The primary sources for this review are the official TypeSafe AI homepage, the official TypeSafe documentation, the creator-authored Jev launch article, and TypeSafe's public SDKs and cookbooks.

Claim Evidence status Primary source
Jev is TypeSafe's first public System One model Official TypeSafe introduction
Jev returns typed decisions rather than generated text Official System One concepts
The primitives are Choice, Score, and Noul Official TypeSafe primitives
Questions over the same state are evaluated independently and in parallel Official TypeSafe introduction
Choice and Score return probabilities and confidence; Noul returns a yes probability Official Choice, Score, Noul
Code should retain control flow and deterministic work Official How to build with TypeSafe
Jev is not a drop-in coding-agent LLM Official Jev with coding agents
Jev is trained with Reinforcement Learning for Calibrated Decisions (RLCD), not RLHF or RLVR Official Jev launch article
Questions should be atomic; decompose multi-factor judgments and combine in code Official TypeSafe introduction
Reported latency of roughly 70–500ms per call, and cost of $0.042 per million input tokens with unmetered output Official, self-reported TypeSafe homepage, Jev launch article
Reported ~193.6x faster and ~444.6x cheaper results on TypeSafe's own workflow evaluations Official, self-reported, vendor methodology Jev launch article, workflow evaluation site
Jev can score retrieved passages and screen prompt injection Official example Classifying RAG passages
Jev can rank a large skill roster and rerank a shortlist Official example Skill suggestion
Jev can screen messages with policy-specific thresholds Official example Guardrails for LLMs
A coding harness should assume no KV cache Creator-derived design exercise Not a Jev API guarantee; use as an engineering constraint
Jev itself stores chunks, reuses model caches, executes tools, or enforces OS permissions Not supported Those remain harness and infrastructure responsibilities

The distinction matters. A typed judgment is not an authorization boundary. A probability is not a policy. A model recommendation is not an executed action.


The Jev programming model in one page

A TypeSafe request contains:

  1. State: the text or structured JSON to evaluate.

  2. Questions: one or more typed judgments over that state.

  3. Model: typically the jev-latest alias or a pinned Jev version.

The three primitives have different meanings:

Primitive Use it when Returned signal
Choice Exactly one option should win from a bounded set Winner, per-option probabilities, confidence
Score The answer lies on ordered, descriptive levels Probability-weighted score, per-level probabilities, confidence
Noul A proposition may be true or false Probability that the answer is yes
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

state = {
    "task": "Fix the expired-coupon checkout bug",
    "candidate_context": {
        "checkout": "Coupon validation and total calculation",
        "email": "Receipt delivery templates",
    },
}

questions = {
    "best_context": Choice(
        instructions="Which context is most relevant to the task?",
        criteria={
            "checkout": "Code that validates coupons and calculates checkout totals",
            "email": "Code that formats and sends receipts",
        },
    ),
    "is_security_sensitive": Noul(
        instructions="Could this task change authorization, payment, or secret-handling behavior?",
    ),
    "execution_risk": Score(
        instructions="How risky would an incorrect automated change be?",
        criteria=[
            "Low: local and easily reversible",
            "Moderate: affects customer-visible behavior but has a clear rollback",
            "High: affects money, authorization, secrets, or production infrastructure",
        ],
    ),
}

with TypeSafeClient() as client:
    response = client.system_one(state=state, questions=questions)

The response is not a patch or an explanation. It is data for code to inspect.

Typed decision loop: state, Jev judgments, harness routing, LLM generation, and policy gates

Method 1: Use Jev as a decision plane, not as the coding model

The first method is conceptual and non-negotiable.

A coding LLM and Jev solve different problems.

A coding LLM:

  • reads open-ended instructions

  • navigates a repository

  • explains hypotheses

  • generates code and prose

  • calls tools over multiple turns

Jev:

  • receives one explicit state

  • evaluates bounded questions

  • returns typed decisions and probabilities

  • does not generate code or explanations

  • does not own the tool loop

TypeSafe's coding-agent guidance says there is no model setting that turns an existing coding agent into a Jev-powered agent. The practical pattern is to use the coding agent to build or operate a harness that calls Jev at selected decision points.

Good Jev responsibilities

  • context relevance

  • skill selection

  • handler routing

  • risk classification

  • trace verification

  • citation support

  • command-policy signals

  • confidence-aware escalation

Responsibilities that stay elsewhere

  • durable storage

  • repository mutation

  • shell execution

  • authentication and authorization

  • deterministic validation

  • test execution

  • rollback

  • final policy enforcement

Best practice: draw the ownership boundary before writing prompts. Every decision should have one owner: deterministic code, Jev, the generative LLM, a human reviewer, or infrastructure policy.


Method 2: Design as if no KV cache exists

The creator's “no KV cache” question is best treated as a design stress test:

If every model call had to pay again for every token, which state would the system truly need?

This is not a claim that Jev manages a coding model's KV cache. It does not. It is a way to expose accidental dependence on ever-growing transcripts.

A cache-independent harness should maintain explicit records such as:

TaskState
  goal
  constraints
  current_plan
  unresolved_questions
  approvals

EvidenceChunk
  chunk_id
  source
  observed_at
  content
  summary
  provenance
  freshness

ActionRecord
  tool
  arguments
  result
  side_effects
  policy_decision
  timestamp

The transcript may still exist for audit, but it is not the only source of truth.

Why this improves engineering

  • Model switches do not erase state.

  • Restarts can reconstruct the task.

  • Context assembly becomes testable.

  • Security constraints are not buried in old dialogue.

  • Retrieval decisions can be audited independently of generation.

Jev's role

Jev can evaluate candidate state fragments, but the harness owns storage and retrieval. TypeSafe's state documentation supports structured JSON and recommends descriptive fields. Its Jev 1.13 limitations warn that irrelevant detail reduces accuracy.

Best practice: store first, retrieve second, judge third, assemble last. Do not use a conversation buffer as the database.


Method 3: Price routing instead of routing by intuition

Model routing can save money, but only when the total path is cheaper and sufficiently reliable.

A route is not free. It may add:

  • another model invocation

  • repeated context input

  • serialization and network latency

  • cache misses

  • inconsistent hidden state

  • another failure and retry surface

The pasted notes include specific routing arithmetic, now cross-checked against a separately compiled reference document, Jev Engineering for Coding Agents (subtitled "The TypeSafe Founder's Blueprint for Building with Jev," dated September 2026). That document states plainly that it is an independent synthesis for study, compiled from design notes by Diogo Almeida, TypeSafe's founder, and that it is not a TypeSafe publication and not affiliated with or endorsed by TypeSafe. It works a concrete example: with illustrative list prices of \(5 input / \)25 output per million tokens for a frontier model and \(3 / \)15 for a mid-tier model, let $X$ be context tokens, $Y$ generated output tokens, and $Z$ additional tokens read during the work (command output, file reads):

$$C_{\text{pure frontier}} = 25Y + 5Z$$

$$C_{\text{frontier} \to \text{mid-tier} \to \text{frontier}} = 3X + 20Y + 8Z$$

Plugging in an illustrative session shape of \(X=0.65\), \(Y=0.12\), \(Z=0.23\) gives 4.15 for staying on the frontier model against 6.19 for the routed path: staying on the frontier model costs roughly two-thirds as much as the route that was supposed to save money. The lesson the notes draw is not that routing is wrong, but that routing priced per token rather than per context rebuild is wrong: routing only pays off once the harness can hand the cheaper model a small, purpose-built context instead of the full transcript. These are illustrative list prices and one session shape from a third-party synthesis, not a general result; recompute the arithmetic from your own observed telemetry before trusting it.

$$C_{route} = \sum_{i=1}^{n} \left(T^{(i)}{in}P^{(i)}{in} + T^{(i)}{out}P^{(i)}{out}\right) + C_{retry} + C_{latency}$$

Where:

  • \(T_{in}\) and \(T_{out}\) are measured tokens

  • \(P_{in}\) and \(P_{out}\) are current provider prices

  • \(C_{retry}\) captures expected retries

  • \(C_{latency}\) expresses the business cost of delay when relevant

Use Jev for bounded routing judgments

TypeSafe documents intent routing: a Choice can identify the request type, a Score can estimate complexity, and code can route to deterministic logic, a specialist LLM, or a human.

route = response.answers["handler"]
complexity = response.answers["complexity"]

if route.confidence < 0.65:
    return human_review(task)

if route.choice == "deterministic":
    return run_static_handler(task)

if complexity.score > 1.5:
    return frontier_coding_model(task)

return economical_coding_model(task)

Best practice: log the complete route, including repeated input tokens, cache status, latency, retries, and outcome quality. Optimize dollars per verified task, not dollars per call.


Method 4: Follow the tokens and measure the real workload

The creator's harness framing argues that coding agents spend much of their budget reading and searching rather than writing code. An independently compiled synthesis of the same design notes publishes an illustrative, input-heavy breakdown of where tokens go in a typical CLI coding-agent session, and cites Microsoft's fastcontext project as independent corroboration that, in GPT-5.4 trajectories, reading and searching account for 56.2% of tool-use turns and 46.5% of the main agent's total tokens:

Phase Illustrative share Note
Reading file contents 30–40% Largest bucket; files reread as context
Searching the codebase 10–18% grep, glob, listings; noisy output
Command output 10–20% Stack traces and logs balloon on failure
System prompt, tool schemas, AGENTS.md 5–12% Fixed overhead paid on every turn
Reasoning and planning 5–15% Higher on hard debugging
Writing and editing code 4–10% Diffs and str_replace edits are compact
Explaining to the user 2–5% Terse by design in CLI agents

Writing code, the task a coding agent exists to perform, is consistently one of the smallest line items. If that pattern holds in your system, the highest-leverage optimization is rarely a better model or diff format; it is smarter, more targeted retrieval. Treat the exact percentages as a prompt to instrument your own system, not as a universal constant.

Measure tokens and elapsed time by phase:

Phase Useful telemetry
Repository discovery files listed, search queries, bytes read, repeated reads
Retrieval candidates, chunks scored, top-k retained, source diversity
Reasoning model, context size, completion size, retries
Tool execution calls, failures, side effects, duration
Validation tests run, regressions found, evaluator results
Recovery repeated failures, revised hypotheses, abandoned branches

What to optimize first

  1. Repeated reading of unchanged files.

  2. Sending irrelevant state to expensive models.

  3. Re-sending large tool schemas.

  4. Serial questions that could be evaluated together.

  5. Background reviewers that repeat retrieval independently.

TypeSafe's parallel questions cookbook reports one documented workload where batching 13 questions over the same large document was 12.2x cheaper and 10.0x faster than 13 sequential calls, with no observed batching effect in that experiment. That is evidence for the principle, not a universal multiplier.

Best practice: establish a cost ledger per completed task and retain the raw counts behind every optimization claim.


Method 5: Score context after the query is known

Blind compaction asks, “What seems important in this transcript?” Query-aware retrieval asks, “What is important for this decision now?”

The second question is better.

A practical context selector can assign each candidate chunk a presentation mode:

  • hide: omit it

  • short: one-line identity or summary

  • long: richer summary with provenance

  • full: original content

Jev does not provide those storage modes automatically. The harness defines them. Jev can supply relevance and risk signals used by the selection policy.

TypeSafe's RAG passage-classification cookbook uses four Noul questions per query-passage pair:

  • Is it relevant?

  • Does it contain usable evidence?

  • Does it contradict the query's premise?

  • Does it attempt prompt injection?

Code then applies thresholds in a fixed order.

if answers["contains_prompt_injection"].noul > 0.70:
    mode = "hide"
elif answers["contradicts_query_premise"].noul > 0.70:
    mode = "full_conflict"
elif answers["is_relevant"].noul < 0.45:
    mode = "hide"
elif answers["contains_answer_evidence"].noul > 0.55:
    mode = "full"
else:
    mode = "short"

The cookbook explicitly warns that a model score is not a security boundary. Retrieved text must still be treated as untrusted.

Query-aware context selection using Jev signals and deterministic presentation policy

Best practice: preserve provenance with every summary, and allow the generator to request the full source when a summary is insufficient.


Method 6: Disclose tools and skills progressively

A large tool catalog creates two separate problems:

  1. Tool definitions consume prompt space.

  2. Similar tools become difficult to distinguish.

The creator's tiered-disclosure idea is supported by TypeSafe's official skill suggestion cookbook.

That cookbook evaluates a two-stage design over 182 skills:

  1. Rank the whole roster using short descriptions.

  2. Rerank a shortlist using full descriptions and instruction excerpts.

  3. Use absolute Noul fit checks to allow “none apply.”

In the published experiment, wrong skill loads fell from 16.8% to 7.3%, and needless loads from 9.8% to 4.0%. These are results for that dataset, harness, and model configuration, not universal guarantees.

A three-tier catalog

Tier Included in context Purpose
Index Name plus one-line description Broad candidate discovery
Schema Arguments, authority, side effects Shortlist evaluation
Full docs Examples, failure modes, prerequisites Final selection and execution

Use relative and absolute judgments together

  • Choice: Which candidate is best among these?

  • One Noul per candidate: Does this candidate actually fit?

A relative winner always exists. Absolute fit checks allow the harness to reject all candidates.

Best practice: never equate “highest probability in the shortlist” with “safe and applicable.” Include a no-match path and validate authority separately.


Method 7: Load instructions conditionally and keep policy durable

Repository instructions should be scoped to the files and actions they govern.

Examples:

  • Editing *.tsx loads UI and accessibility guidance.

  • Touching billing/ loads monetary invariants and approval requirements.

  • Modifying infrastructure loads environment and rollback policy.

  • Opening a migration file loads data-loss and idempotency checks.

This is a harness design pattern, not a built-in Jev instruction loader.

The harness should maintain an explicit rule index:

instruction_rules:
  - match: "**/*.tsx"
    include:
      - policies/frontend-accessibility.md
      - policies/design-system.md
  - match: "billing/**"
    include:
      - policies/money-invariants.md
    require_approval: true
  - match: "infra/**"
    include:
      - policies/infrastructure-change.md
    require_approval: true

Jev can help judge ambiguous applicability, but deterministic patterns should remain deterministic.

Why this is safer than transcript-only instructions

  • Compaction cannot silently delete policy.

  • Restarts can reconstruct active rules.

  • The audit log can name which policy applied.

  • Reviewers can test rule activation without invoking a model.

TypeSafe's building guide recommends structured state, narrow questions, and code-owned control flow. Jev's documented jaggedness also warns against excessive indirection and irrelevant state.

Best practice: use code to load rules, Jev to resolve semantic ambiguity, and infrastructure to enforce the final boundary.


Method 8: Route by trust and impact, not only by difficulty

Difficulty is only one routing dimension. Enterprise systems also need to consider:

  • data sensitivity

  • action reversibility

  • financial impact

  • authorization scope

  • source trust

  • required auditability

A simple documentation lookup may be intellectually easy but confidential. A difficult public-code explanation may be safe for a cheaper external model. The route should reflect both capability and trust.

A two-axis route

Capability requirement: low -> high
Trust requirement:      public -> restricted -> privileged

Jev can classify intent and score risk. Code chooses the permitted destination.

if data_classification == "privileged":
    handler = approved_private_model
elif risk.score > 1.5 or risk.confidence < 0.6:
    handler = frontier_model_with_review
elif task_type.choice == "documentation":
    handler = economical_model
else:
    handler = standard_coding_model

TypeSafe's confidence-routing guidance stresses that thresholds should scale with consequences. Showing the wrong screen can tolerate a lower threshold than approving a transfer.

An independently compiled synthesis of the design notes frames this as routing by data sensitivity rather than difficulty alone, attaching an eligible-model policy to the kind of file a subtask is likely to touch:

Files likely touched Policy Eligible models
Public docs, open-source dependencies open any, cheapest first
Application code standard vetted providers
Secrets, environment, infrastructure config restricted first-party frontier only
Proprietary research code custom excludes named vendors

The underlying point generalizes beyond that specific table: once routing is policy-driven rather than ad hoc, a preference to avoid a given vendor for a class of files becomes configuration the harness enforces automatically, rather than discipline every engineer must remember.

Best practice: permission checks must occur before model routing. Confidence can tighten policy; it must never grant authority the caller does not already possess.


Method 9: Share retrieval, then fan out independent review

Background reviewers are useful when they add different judgments, not when each one repeats the same expensive discovery pass.

A better structure is:

  1. Retrieve and normalize evidence once.

  2. Freeze an evidence-pack version.

  3. Ask multiple independent questions over that same state.

  4. Send only tasks needing generation to LLM reviewers.

  5. Merge structured results by policy.

TypeSafe calls this speculative fan-out: ask independent questions together, including questions that may only matter on one branch, and let code ignore irrelevant answers.

Speculative fan-out over one shared evidence pack

Good fan-out questions for a code patch

  • Does the patch address the reproduced failure?

  • Does it modify unrelated behavior?

  • Does it touch an authorization boundary?

  • Does it add a dependency?

  • Does it contradict repository instructions?

  • Does it need human approval?

Each question should be narrow and independent. TypeSafe explicitly warns that broad questions hide multiple judgments.

Best practice: immutable evidence packs make reviewer disagreements diagnosable. If two reviewers saw different state, their disagreement is not meaningful.


Method 10: Gate every consequential command

A model should not be the final authority over shell commands, deployments, data deletion, or production mutation.

A robust gate has three outcomes:

  • allow: low-risk and within existing authority

  • ask: requires human approval or clarification

  • deny: prohibited regardless of model confidence

Jev can contribute semantic signals:

  • Does the script delete data?

  • Does it access secrets?

  • Does it modify production infrastructure?

  • Does it match the stated task?

  • Is a rollback path present?

Code and infrastructure must make the final decision.

assessment = client.system_one(
    state={
        "task": task,
        "command": command,
        "working_directory": cwd,
        "policy": active_policy,
    },
    questions={
        "matches_task": Noul(
            instructions="Does the command directly support the stated task?"
        ),
        "is_destructive": Noul(
            instructions="Could the command delete, overwrite, or irreversibly mutate data?"
        ),
        "risk": Score(
            instructions="How much damage could an incorrect execution cause?",
            criteria=[
                "Local and immediately reversible",
                "Limited impact with a tested rollback",
                "Production, security, financial, or irreversible impact",
            ],
        ),
    },
)

answers = assessment.answers

if hard_policy_denies(command, identity, cwd):
    decision = "deny"
elif answers["is_destructive"].noul > 0.45:
    decision = "ask"
elif answers["matches_task"].noul < 0.70:
    decision = "ask"
elif answers["risk"].score > 1.25:
    decision = "ask"
else:
    decision = "allow"

TypeSafe's LLM guardrails cookbook follows the same division: Jev supplies assessments; application policy maps them to pass, review, block, or support. The same probabilities can produce different actions under different named policies.

An independently compiled synthesis of the design notes expresses the deterministic half of that gate as a small policy language, with Jev's semantic signals feeding the conditions rather than replacing them:

policy "exec":
  deny   if command touches ~/.ssh or .env*
  deny   if script contents contain network egress
         and task.scope != "deploy"
  ask    if command writes outside repo root
  allow  if command in read_only_set
  allow  if tests/ and exit code is expected

The same source recommends inspecting script contents, not only the command name, before execution, for example reading a Python or shell file before running it rather than approving the executable alone. This is a design pattern to adapt, not an official Jev policy DSL.

What a command gate must log

  • requested command and normalized script

  • requesting identity

  • active task and policy version

  • Jev model version

  • raw probabilities and scores

  • deterministic policy checks

  • final decision

  • approving human, if any

  • execution result and side effects

Best practice: inspect the full script and arguments, not only the tool name. terminal is not inherently dangerous; terminal("Remove-Item -Recurse ...") may be.


The complete harness loop

The ten methods combine into a runtime in which state, decisions, execution, and verification remain separate.

Complete Jev-assisted coding-agent harness lifecycle

Reference pseudocode

while not task.finished:
    state = store.load(task.id)
    candidates = retrieve(state.goal, state.repo_snapshot)

    context_signals = jev_score_context(state.goal, candidates)
    context = assemble_context(candidates, context_signals, token_budget)

    route = jev_route(state, context, available_handlers)
    proposal = coding_llm.run(state.goal, context, route)

    for action in proposal.actions:
        policy = deterministic_policy(action, identity, environment)
        semantic = jev_assess_action(state, action, policy)
        decision = combine_policy(policy, semantic)

        if decision == "deny":
            store.record_denial(task.id, action, policy, semantic)
            continue
        if decision == "ask":
            if not human_approval(action, policy, semantic):
                continue

        observation = execute_in_sandbox(action)
        store.append_observation(task.id, observation)

    validation = run_targeted_checks(task)
    verification = jev_verify_trace(store.trace(task.id), validation)
    task.finished = completion_policy(validation, verification)

This pseudocode is an architectural example, not an official Jev harness SDK.


Failure modes and design limits

TypeSafe publishes a Jev 1.13 jaggedness page. That candor is important because typed output is not the same as semantic correctness.

Documented limitations include:

  • literal interpretation

  • unreliable counting and arithmetic

  • weak date ordering and comparison

  • reduced accuracy with indirection

  • context rot from irrelevant state

  • susceptibility to adversarial content

  • confusion from contradictory instructions and criteria

  • no guaranteed arithmetic invariants between separately asked questions

  • no text generation

Engineering responses

Limitation Harness response
Math and counting Compute deterministically in code
Dates Extract components, then compare in code
Irrelevant state Retrieve and filter before calling Jev
Adversarial content Treat state as untrusted; test and enforce policy outside the model
Literal reading Write explicit boundaries and examples
Structural inconsistency Do not assume separate probabilities obey identities
Generation Use an LLM for text and code generation
Confidence Validate thresholds on domain data; do not treat confidence as truth

Type safety guarantees that the answer matches the declared shape. It does not guarantee that the judgment is correct.


How to evaluate a Jev-assisted coding harness

A production evaluation should compare systems, not isolated model calls.

Quality metrics

  • relevant context retained

  • irrelevant context removed

  • correct skill or tool suggested

  • false command allows

  • false command blocks

  • task completion rate

  • regression rate

  • human-review precision and recall

Cost and performance metrics

  • total tokens across all models

  • Jev input and output usage

  • number of model round trips

  • cache hit rate

  • end-to-end latency

  • cost per verified task

Reliability metrics

  • variance across repeated runs

  • recovery after tool failure

  • restart recovery from durable state

  • policy adherence after context compaction

  • performance on adversarial retrieved text

  • calibration of confidence against observed accuracy

An ablation plan

Run the same task set with:

  1. Baseline coding agent.

  2. Agent plus query-aware context scoring.

  3. Agent plus progressive tool disclosure.

  4. Agent plus confidence-aware routing.

  5. Agent plus command gating.

  6. Complete harness.

This shows which components add value and which only add complexity.


What the ten methods solve

Inherited problem Engineering response
Repeated context processing Durable state and bounded context assembly
Blind model routing Measured, confidence-aware routing
Tool-schema bloat Progressive disclosure and shortlist reranking
Blind compaction Query-aware chunk judgment
Implicit sub-agent state Versioned evidence packs
Lost state on restart Explicit persisted task records
Unsafe execution Deterministic policy plus semantic command assessment
Reviewer duplication Shared retrieval and parallel independent judgments
Unclear uncertainty Probabilities, confidence, and review paths
Opaque workflow Typed decisions and auditable code-owned control flow

Current public applications and experiments

Jev's public ecosystem is still early. The examples below are applications that could be verified through TypeSafe's official documentation, cookbooks, evaluation site, launch material, or public repositories as of September 2026. They should not be read as a list of independently confirmed customer production deployments.

Interactive and runnable demos

Application What Jev does Evidence level Source
Smart-home assistant Classifies request category, room, device, and action; falls back to an LLM for conversation or compound-request splitting Official interactive demo Smart-home assistant
Natural-language trading interface Selects one of ten typed functions and fills bounded arguments from natural language Reproducible official cookbook Function calling
Jev Playground Evaluates user-provided state with Choice, Score, and Noul questions Official hosted application TypeSafe Playground

The smart-home demo is a particularly clear example of the intended architecture. Jev handles bounded decisions in parallel; code filters the answers; an LLM is invoked only when the request requires free-form generation.

Agent and developer-tooling applications

Application What Jev does Published result or behavior Source
Hermes skill suggestion Ranks 182 skills, reranks a shortlist, and can recommend no skill In the published 488-request experiment, wrong loads fell from 16.8% to 7.3% and needless loads from 9.8% to 4.0% Skill suggestion
TypeSafe agent skill Gives Claude Code, Codex, and other coding agents current TypeSafe integration guidance Public installable skill; it helps agents write TypeSafe integrations but does not replace their LLM Agent skill repository
System One adapter Runs the same TypeSafe-shaped questions through OpenAI-compatible or Anthropic LLM APIs Public comparison and compatibility tool with retries and diagnostics System One adapter

The Hermes experiment is the closest verified public example to the coding-harness design in this article. Jev does not load the skill or run the agent. It supplies a bounded suggestion that the existing agent may follow or ignore.

Retrieval, verification, and safety applications

Application Jev's role Source
RAG passage gating Scores relevance, usable evidence, premise contradiction, and prompt injection before generation Classifying RAG passages
Citation verification Classifies whether source context supports, contradicts, or says nothing about a claim Double-checking citations
LLM guardrails Screens inputs and outputs for jailbreaks, harmful requests, medical advice, self-harm, and severity Guardrails for LLMs
Legal retrieval reranking Reranks BM25 candidates for legal queries Re-ranking cookbook
Knowledge-graph entity alignment Judges whether candidate records represent the same entity and surfaces disagreeing fields Entity alignment
Structured extraction cascade Verifies extracted fields and escalates difficult cases to a reasoning model SDE cascade

These are worked examples, not generic claims. Their thresholds, model versions, datasets, and reported results belong to those specific experiments.

Published workflow evaluations

TypeSafe's workflow evaluation site publishes four larger policy-shaped applications:

  1. Security incidents: decide whether to close an alert, send it to an analyst, or contain a machine.

  2. Agent trace observability: inspect a completed support-agent trace and decide whether human review is needed and how urgently.

  3. Invoice processing: decide whether to pay, hold, or return a vendor invoice using the invoice, purchase order, and delivery state.

  4. Customer service: decide what a support system should do next from the conversation and account state.

The evaluation assumes the workflow itself is correct and compares model judgments inside that fixed compute graph. It is evidence about structured workflow performance under TypeSafe's published methodology, not independent proof that the workflows are production deployments.

Launch demonstrations

The creator-authored Jev launch article also describes:

  • a Doom bot making rapid decisions from structured textual game state

  • Wikipedia racing, where Jev repeatedly chooses links from large candidate sets

These demonstrations illustrate latency, bounded choice, and repeated decision-making. TypeSafe notes important caveats: the Doom demo uses structured text rather than vision, and the Wikipedia-racing comparison settings affect the observed speedups.

What could be verified from X

The official TypeSafe AI X account is public and identifies the company as an AI lab building intelligence beyond chat. However, unauthenticated access during this review exposed the profile but not searchable post bodies. X search pages required login, so this article does not attribute any application, performance result, or third-party deployment to an X post that could not be read and linked directly.

No independently verifiable third-party production deployment was identified in the public sources reviewed here. That does not mean none exist; it means the evidence currently available for this article is strongest for TypeSafe's own runnable demos, cookbooks, SDKs, repositories, and published evaluations.


Conclusion

The most important idea in Jev engineering is not that a smaller model should replace a larger one.

It is that generation and decision-making do not need to be the same operation.

A generative coding model is useful because software work is open-ended. It must read unfamiliar code, construct hypotheses, explain failures, and create patches. But many decisions inside that workflow are bounded:

  • choose one tool

  • score one risk dimension

  • judge whether one condition holds

  • rank candidate context

  • decide whether uncertainty requires escalation

Jev is designed for that bounded layer.

The architecture becomes easier to reason about when:

  • state is explicit

  • evidence has provenance

  • questions are atomic

  • probabilities remain visible

  • code owns thresholds and control flow

  • infrastructure owns authorization

  • the LLM owns generation

  • humans own consequential approvals

That is the deeper lesson behind the Jev harness idea:

Do not ask one model to remember, decide, generate, authorize, execute, and verify everything. Give each responsibility to the component that can make it explicit, testable, and governable.

Jev's opportunity is therefore architectural. It offers a machine-oriented decision interface at points where agent systems often rely on hidden prompt logic: context selection, routing, tool choice, verification, risk assessment, and escalation. Its typed outputs make those decisions easier to inspect and compose, while probabilities make uncertainty available to policy code.

Its limits are equally important. Jev does not make authorization unnecessary, does not turn model confidence into truth, does not remove prompt-injection risk, and does not replace a generative coding model. Its current documented jagged edges include literal reading, context rot, numeric weakness, date-comparison weakness, adversarial influence, and lack of text generation.

The strongest implementation path is incremental:

  1. Pick one bounded decision already causing cost, latency, or reliability problems.

  2. Define the state and one atomic typed question.

  3. Keep deterministic rules and side effects in code.

  4. Validate probabilities and thresholds on representative data.

  5. Add a human or reasoning-model route for uncertain cases.

  6. Measure the complete workflow, not only the Jev call.

If that first decision improves the system, add another. The goal is not to place Jev everywhere. The goal is to replace opaque, repeated, free-form judgment with small decision interfaces wherever doing so makes the harness more reliable, economical, and governable.


Official resources

5 views