Jev Engineering for Coding-Agent Harnesses
TL;DR
Jev is TypeSafe AI's "System One" model — not a chat model, not a replacement for the LLM in a coding agent. It returns bounded, typed judgments (
Choice,Score,Noul) that software can consume directly.The useful architecture is: code owns the workflow, the coding LLM generates, and Jev supplies narrow semantic judgments at selected control points (relevance filtering, tool routing, prompt-injection detection, policy gating, escalation to a human).
This article synthesizes ten harness-engineering methods built around that division of labor, sourced from Diogo Almeida's published design notes and TypeSafe's documentation.
The goal is a coding-agent harness that is cheaper, more explainable, and more governable than "one giant LLM prompt with a loop."
Introduction to Jev
Coding agents are usually described as if the model were the whole product. Give a frontier language model a long prompt, expose a terminal and a few file tools, add a loop, and hope that the model keeps the task, constraints, evidence, and safety policy straight.
That design can work. It can also become expensive, opaque, and difficult to govern.
Jev suggests a different division of labor.
This entire article is based on the work of Diogo Almeida, founder of TypeSafe and creator of Jev. The ten harness-engineering methods below synthesize his published design notes on building coding-agent harnesses around Jev, including the "no KV cache" thought experiment, the routing arithmetic, the token-share findings, and the tiered-disclosure and command-gating patterns. Credit for that underlying thinking belongs to him; this article organizes, cross-checks, and sources it against TypeSafe's official documentation.
Jev is TypeSafe AI's first public System One model. Jev is not a chat model, and it is not a replacement for the LLM inside Claude Code, Cursor, GitHub Copilot, Codex, or another coding agent. TypeSafe's own guide is explicit: Jev does not generate text, write code, hold a conversation, or choose its own next action. It evaluates structured state against typed questions and returns bounded decisions that software can consume directly.
That makes Jev potentially useful inside the harness around a coding LLM.
The LLM can still investigate, explain, and write code. The harness can retain state, execute tools, enforce permissions, and verify outcomes. Jev can supply narrow semantic judgments where deterministic code is too brittle:
Which context chunks are relevant to this task?
Which skill or tool best matches the request?
Does a retrieved passage contain prompt injection?
Does a proposed command conflict with policy?
Which handler should receive the next step?
Is the evidence strong enough to act, or should the system ask for review?
The useful architecture is not “Jev replaces the coding agent.” It is “code owns the workflow, the coding LLM generates, and Jev supplies typed judgments at selected control points.”
Why Jev exists
TypeSafe's thesis is that chat models and machine-consumed decisions are different workloads. Chat models are optimized to generate useful strings for people. Jev is designed to return bounded values, probability distributions, and uncertainty signals that software can inspect without parsing prose.
TypeSafe calls Jev a System One model, borrowing the name from the fast, focused side of the System 1/System 2 distinction. In this programming model:
application state is passed explicitly as text or structured JSON
questions declare the judgment and its valid answer space
Choiceselects among bounded alternativesScoreplaces state on ordered, descriptive levelsNoulreturns the probability that a proposition is trueindependent questions over the same state can be evaluated together
application code decides how probabilities become actions
This makes Jev closer to a semantic decision primitive than a conversational assistant. It can help code answer questions that ordinary if statements cannot answer from raw text, while preserving a bounded output contract.
How Jev is trained, and what TypeSafe claims about cost and speed
TypeSafe's launch article describes a training method it calls Reinforcement Learning for Calibrated Decisions (RLCD), contrasted with the RLHF and RLVR methods used to train conversational LLMs. Where RLHF optimizes for responses human raters prefer, TypeSafe states RLCD optimizes for epistemically honest, calibrated probabilities on bounded System One tasks. Jev also uses a parallel sampler: instead of generating one token at a time, it evaluates all questions against a state in a single pass.
TypeSafe's own published figures, current as of the September 2026 launch, include:
Latency: an end-to-end response time of roughly 70–500ms per call, versus 3–329 seconds TypeSafe attributes to frontier LLMs on comparable tasks.
Cost: input tokens priced at $0.042 per million tokens (about $42 per billion tokens), with output pricing described as unmetered because Jev does not generate free text.
Workflow evaluations: TypeSafe reports up to roughly 193.6x faster and 444.6x cheaper results on its own published "workflow eval" methodology, which compares Jev's typed decisions against an average of large reference LLMs across four workflows (security incidents, agent-trace observability, invoice processing, and customer service).
Type safety: TypeSafe states Jev cannot produce a schema-mismatched output, because valid outputs are defined in advance and enforced structurally, rather than validated after the fact the way free-text LLM output is.
These are TypeSafe's own self-reported claims, published alongside methodology notes and caveats on the same page (for example, that its workflow evaluation harness and prompts were authored by TypeSafe's own team, and that its comparison baseline is an average of two specific frontier models). Treat them as vendor-published evidence to verify against your own workload, not as independently audited benchmark results.
TypeSafe's documentation also gives explicit guidance on question design: keep each Choice, Score, or Noul question atomic, meaning it should ask one well-scoped judgment, similar to a determination a knowledgeable person could make in a few seconds given the right context. If a judgment would require weighing multiple independent factors, the guidance is to decompose it into separate questions and combine the results in code, rather than asking Jev to do the weighing implicitly. This directly supports the harness-side composition patterns developed later in this article.
Why it matters for coding-agent harnesses
A coding agent contains many decisions that are narrower than code generation itself: select context, rank skills, detect risky content, classify intent, score a patch, or decide whether uncertainty requires review. Those decisions can be separated from the generative model and made visible to the harness.
The result is not a fully autonomous Jev coding agent. It is a hybrid system:
| Component | Primary responsibility |
|---|---|
| Coding LLM | Open-ended investigation, explanation, and code generation |
| Jev | Narrow typed judgments with probabilities |
| Harness code | State, thresholds, routing, retries, and composition |
| Policy engine | Deterministic authorization and prohibited actions |
| Tools and sandbox | Scoped execution and side effects |
| Evaluators | Tests, regression checks, and outcome verification |
| Human reviewer | Approval for consequential or uncertain actions |
Ten methods for building explicit, typed, confidence-aware agent systems
This article develops ten engineering methods for that architecture. Some originate in the creator's harness framing. Others are supported directly by TypeSafe's documentation and cookbooks. Where the evidence differs, the distinction is stated explicitly.
Source boundary: what is official and what is architectural interpretation
Jev is new, and its public material is evolving quickly. A responsible design should separate confirmed product behavior from proposed harness architecture.
The primary sources for this review are the official TypeSafe AI homepage, the official TypeSafe documentation, the creator-authored Jev launch article, and TypeSafe's public SDKs and cookbooks.
| Claim | Evidence status | Primary source |
|---|---|---|
| Jev is TypeSafe's first public System One model | Official | TypeSafe introduction |
| Jev returns typed decisions rather than generated text | Official | System One concepts |
| The primitives are Choice, Score, and Noul | Official | TypeSafe primitives |
| Questions over the same state are evaluated independently and in parallel | Official | TypeSafe introduction |
| Choice and Score return probabilities and confidence; Noul returns a yes probability | Official | Choice, Score, Noul |
| Code should retain control flow and deterministic work | Official | How to build with TypeSafe |
| Jev is not a drop-in coding-agent LLM | Official | Jev with coding agents |
| Jev is trained with Reinforcement Learning for Calibrated Decisions (RLCD), not RLHF or RLVR | Official | Jev launch article |
| Questions should be atomic; decompose multi-factor judgments and combine in code | Official | TypeSafe introduction |
| Reported latency of roughly 70–500ms per call, and cost of $0.042 per million input tokens with unmetered output | Official, self-reported | TypeSafe homepage, Jev launch article |
| Reported ~193.6x faster and ~444.6x cheaper results on TypeSafe's own workflow evaluations | Official, self-reported, vendor methodology | Jev launch article, workflow evaluation site |
| Jev can score retrieved passages and screen prompt injection | Official example | Classifying RAG passages |
| Jev can rank a large skill roster and rerank a shortlist | Official example | Skill suggestion |
| Jev can screen messages with policy-specific thresholds | Official example | Guardrails for LLMs |
| A coding harness should assume no KV cache | Creator-derived design exercise | Not a Jev API guarantee; use as an engineering constraint |
| Jev itself stores chunks, reuses model caches, executes tools, or enforces OS permissions | Not supported | Those remain harness and infrastructure responsibilities |
The distinction matters. A typed judgment is not an authorization boundary. A probability is not a policy. A model recommendation is not an executed action.
The Jev programming model in one page
A TypeSafe request contains:
State: the text or structured JSON to evaluate.
Questions: one or more typed judgments over that state.
Model: typically the
jev-latestalias or a pinned Jev version.
The three primitives have different meanings:
| Primitive | Use it when | Returned signal |
|---|---|---|
| Choice | Exactly one option should win from a bounded set | Winner, per-option probabilities, confidence |
| Score | The answer lies on ordered, descriptive levels | Probability-weighted score, per-level probabilities, confidence |
| Noul | A proposition may be true or false | Probability that the answer is yes |
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
state = {
"task": "Fix the expired-coupon checkout bug",
"candidate_context": {
"checkout": "Coupon validation and total calculation",
"email": "Receipt delivery templates",
},
}
questions = {
"best_context": Choice(
instructions="Which context is most relevant to the task?",
criteria={
"checkout": "Code that validates coupons and calculates checkout totals",
"email": "Code that formats and sends receipts",
},
),
"is_security_sensitive": Noul(
instructions="Could this task change authorization, payment, or secret-handling behavior?",
),
"execution_risk": Score(
instructions="How risky would an incorrect automated change be?",
criteria=[
"Low: local and easily reversible",
"Moderate: affects customer-visible behavior but has a clear rollback",
"High: affects money, authorization, secrets, or production infrastructure",
],
),
}
with TypeSafeClient() as client:
response = client.system_one(state=state, questions=questions)
The response is not a patch or an explanation. It is data for code to inspect.
Method 1: Use Jev as a decision plane, not as the coding model
The first method is conceptual and non-negotiable.
A coding LLM and Jev solve different problems.
A coding LLM:
reads open-ended instructions
navigates a repository
explains hypotheses
generates code and prose
calls tools over multiple turns
Jev:
receives one explicit state
evaluates bounded questions
returns typed decisions and probabilities
does not generate code or explanations
does not own the tool loop
TypeSafe's coding-agent guidance says there is no model setting that turns an existing coding agent into a Jev-powered agent. The practical pattern is to use the coding agent to build or operate a harness that calls Jev at selected decision points.
Good Jev responsibilities
context relevance
skill selection
handler routing
risk classification
trace verification
citation support
command-policy signals
confidence-aware escalation
Responsibilities that stay elsewhere
durable storage
repository mutation
shell execution
authentication and authorization
deterministic validation
test execution
rollback
final policy enforcement
Best practice: draw the ownership boundary before writing prompts. Every decision should have one owner: deterministic code, Jev, the generative LLM, a human reviewer, or infrastructure policy.
Method 2: Design as if no KV cache exists
The creator's “no KV cache” question is best treated as a design stress test:
If every model call had to pay again for every token, which state would the system truly need?
This is not a claim that Jev manages a coding model's KV cache. It does not. It is a way to expose accidental dependence on ever-growing transcripts.
A cache-independent harness should maintain explicit records such as:
TaskState
goal
constraints
current_plan
unresolved_questions
approvals
EvidenceChunk
chunk_id
source
observed_at
content
summary
provenance
freshness
ActionRecord
tool
arguments
result
side_effects
policy_decision
timestamp
The transcript may still exist for audit, but it is not the only source of truth.
Why this improves engineering
Model switches do not erase state.
Restarts can reconstruct the task.
Context assembly becomes testable.
Security constraints are not buried in old dialogue.
Retrieval decisions can be audited independently of generation.
Jev's role
Jev can evaluate candidate state fragments, but the harness owns storage and retrieval. TypeSafe's state documentation supports structured JSON and recommends descriptive fields. Its Jev 1.13 limitations warn that irrelevant detail reduces accuracy.
Best practice: store first, retrieve second, judge third, assemble last. Do not use a conversation buffer as the database.
Method 3: Price routing instead of routing by intuition
Model routing can save money, but only when the total path is cheaper and sufficiently reliable.
A route is not free. It may add:
another model invocation
repeated context input
serialization and network latency
cache misses
inconsistent hidden state
another failure and retry surface
The pasted notes include specific routing arithmetic, now cross-checked against a separately compiled reference document, Jev Engineering for Coding Agents (subtitled "The TypeSafe Founder's Blueprint for Building with Jev," dated September 2026). That document states plainly that it is an independent synthesis for study, compiled from design notes by Diogo Almeida, TypeSafe's founder, and that it is not a TypeSafe publication and not affiliated with or endorsed by TypeSafe. It works a concrete example: with illustrative list prices of \(5 input / \)25 output per million tokens for a frontier model and \(3 / \)15 for a mid-tier model, let $X$ be context tokens, $Y$ generated output tokens, and $Z$ additional tokens read during the work (command output, file reads):
$$C_{\text{pure frontier}} = 25Y + 5Z$$
$$C_{\text{frontier} \to \text{mid-tier} \to \text{frontier}} = 3X + 20Y + 8Z$$
Plugging in an illustrative session shape of \(X=0.65\), \(Y=0.12\), \(Z=0.23\) gives 4.15 for staying on the frontier model against 6.19 for the routed path: staying on the frontier model costs roughly two-thirds as much as the route that was supposed to save money. The lesson the notes draw is not that routing is wrong, but that routing priced per token rather than per context rebuild is wrong: routing only pays off once the harness can hand the cheaper model a small, purpose-built context instead of the full transcript. These are illustrative list prices and one session shape from a third-party synthesis, not a general result; recompute the arithmetic from your own observed telemetry before trusting it.
$$C_{route} = \sum_{i=1}^{n} \left(T^{(i)}{in}P^{(i)}{in} + T^{(i)}{out}P^{(i)}{out}\right) + C_{retry} + C_{latency}$$
Where:
\(T_{in}\) and \(T_{out}\) are measured tokens
\(P_{in}\) and \(P_{out}\) are current provider prices
\(C_{retry}\) captures expected retries
\(C_{latency}\) expresses the business cost of delay when relevant
Use Jev for bounded routing judgments
TypeSafe documents intent routing: a Choice can identify the request type, a Score can estimate complexity, and code can route to deterministic logic, a specialist LLM, or a human.
route = response.answers["handler"]
complexity = response.answers["complexity"]
if route.confidence < 0.65:
return human_review(task)
if route.choice == "deterministic":
return run_static_handler(task)
if complexity.score > 1.5:
return frontier_coding_model(task)
return economical_coding_model(task)
Best practice: log the complete route, including repeated input tokens, cache status, latency, retries, and outcome quality. Optimize dollars per verified task, not dollars per call.
Method 4: Follow the tokens and measure the real workload
The creator's harness framing argues that coding agents spend much of their budget reading and searching rather than writing code. An independently compiled synthesis of the same design notes publishes an illustrative, input-heavy breakdown of where tokens go in a typical CLI coding-agent session, and cites Microsoft's fastcontext project as independent corroboration that, in GPT-5.4 trajectories, reading and searching account for 56.2% of tool-use turns and 46.5% of the main agent's total tokens:
| Phase | Illustrative share | Note |
|---|---|---|
| Reading file contents | 30–40% | Largest bucket; files reread as context |
| Searching the codebase | 10–18% | grep, glob, listings; noisy output |
| Command output | 10–20% | Stack traces and logs balloon on failure |
| System prompt, tool schemas, AGENTS.md | 5–12% | Fixed overhead paid on every turn |
| Reasoning and planning | 5–15% | Higher on hard debugging |
| Writing and editing code | 4–10% | Diffs and str_replace edits are compact |
| Explaining to the user | 2–5% | Terse by design in CLI agents |
Writing code, the task a coding agent exists to perform, is consistently one of the smallest line items. If that pattern holds in your system, the highest-leverage optimization is rarely a better model or diff format; it is smarter, more targeted retrieval. Treat the exact percentages as a prompt to instrument your own system, not as a universal constant.
Measure tokens and elapsed time by phase:
| Phase | Useful telemetry |
|---|---|
| Repository discovery | files listed, search queries, bytes read, repeated reads |
| Retrieval | candidates, chunks scored, top-k retained, source diversity |
| Reasoning | model, context size, completion size, retries |
| Tool execution | calls, failures, side effects, duration |
| Validation | tests run, regressions found, evaluator results |
| Recovery | repeated failures, revised hypotheses, abandoned branches |
What to optimize first
Repeated reading of unchanged files.
Sending irrelevant state to expensive models.
Re-sending large tool schemas.
Serial questions that could be evaluated together.
Background reviewers that repeat retrieval independently.
TypeSafe's parallel questions cookbook reports one documented workload where batching 13 questions over the same large document was 12.2x cheaper and 10.0x faster than 13 sequential calls, with no observed batching effect in that experiment. That is evidence for the principle, not a universal multiplier.
Best practice: establish a cost ledger per completed task and retain the raw counts behind every optimization claim.
Method 5: Score context after the query is known
Blind compaction asks, “What seems important in this transcript?” Query-aware retrieval asks, “What is important for this decision now?”
The second question is better.
A practical context selector can assign each candidate chunk a presentation mode:
hide: omit it
short: one-line identity or summary
long: richer summary with provenance
full: original content
Jev does not provide those storage modes automatically. The harness defines them. Jev can supply relevance and risk signals used by the selection policy.
TypeSafe's RAG passage-classification cookbook uses four Noul questions per query-passage pair:
Is it relevant?
Does it contain usable evidence?
Does it contradict the query's premise?
Does it attempt prompt injection?
Code then applies thresholds in a fixed order.
if answers["contains_prompt_injection"].noul > 0.70:
mode = "hide"
elif answers["contradicts_query_premise"].noul > 0.70:
mode = "full_conflict"
elif answers["is_relevant"].noul < 0.45:
mode = "hide"
elif answers["contains_answer_evidence"].noul > 0.55:
mode = "full"
else:
mode = "short"
The cookbook explicitly warns that a model score is not a security boundary. Retrieved text must still be treated as untrusted.
Best practice: preserve provenance with every summary, and allow the generator to request the full source when a summary is insufficient.
Method 6: Disclose tools and skills progressively
A large tool catalog creates two separate problems:
Tool definitions consume prompt space.
Similar tools become difficult to distinguish.
The creator's tiered-disclosure idea is supported by TypeSafe's official skill suggestion cookbook.
That cookbook evaluates a two-stage design over 182 skills:
Rank the whole roster using short descriptions.
Rerank a shortlist using full descriptions and instruction excerpts.
Use absolute
Noulfit checks to allow “none apply.”
In the published experiment, wrong skill loads fell from 16.8% to 7.3%, and needless loads from 9.8% to 4.0%. These are results for that dataset, harness, and model configuration, not universal guarantees.
A three-tier catalog
| Tier | Included in context | Purpose |
|---|---|---|
| Index | Name plus one-line description | Broad candidate discovery |
| Schema | Arguments, authority, side effects | Shortlist evaluation |
| Full docs | Examples, failure modes, prerequisites | Final selection and execution |
Use relative and absolute judgments together
Choice: Which candidate is best among these?One
Noulper candidate: Does this candidate actually fit?
A relative winner always exists. Absolute fit checks allow the harness to reject all candidates.
Best practice: never equate “highest probability in the shortlist” with “safe and applicable.” Include a no-match path and validate authority separately.
Method 7: Load instructions conditionally and keep policy durable
Repository instructions should be scoped to the files and actions they govern.
Examples:
Editing
*.tsxloads UI and accessibility guidance.Touching
billing/loads monetary invariants and approval requirements.Modifying infrastructure loads environment and rollback policy.
Opening a migration file loads data-loss and idempotency checks.
This is a harness design pattern, not a built-in Jev instruction loader.
The harness should maintain an explicit rule index:
instruction_rules:
- match: "**/*.tsx"
include:
- policies/frontend-accessibility.md
- policies/design-system.md
- match: "billing/**"
include:
- policies/money-invariants.md
require_approval: true
- match: "infra/**"
include:
- policies/infrastructure-change.md
require_approval: true
Jev can help judge ambiguous applicability, but deterministic patterns should remain deterministic.
Why this is safer than transcript-only instructions
Compaction cannot silently delete policy.
Restarts can reconstruct active rules.
The audit log can name which policy applied.
Reviewers can test rule activation without invoking a model.
TypeSafe's building guide recommends structured state, narrow questions, and code-owned control flow. Jev's documented jaggedness also warns against excessive indirection and irrelevant state.
Best practice: use code to load rules, Jev to resolve semantic ambiguity, and infrastructure to enforce the final boundary.
Method 8: Route by trust and impact, not only by difficulty
Difficulty is only one routing dimension. Enterprise systems also need to consider:
data sensitivity
action reversibility
financial impact
authorization scope
source trust
required auditability
A simple documentation lookup may be intellectually easy but confidential. A difficult public-code explanation may be safe for a cheaper external model. The route should reflect both capability and trust.
A two-axis route
Capability requirement: low -> high
Trust requirement: public -> restricted -> privileged
Jev can classify intent and score risk. Code chooses the permitted destination.
if data_classification == "privileged":
handler = approved_private_model
elif risk.score > 1.5 or risk.confidence < 0.6:
handler = frontier_model_with_review
elif task_type.choice == "documentation":
handler = economical_model
else:
handler = standard_coding_model
TypeSafe's confidence-routing guidance stresses that thresholds should scale with consequences. Showing the wrong screen can tolerate a lower threshold than approving a transfer.
An independently compiled synthesis of the design notes frames this as routing by data sensitivity rather than difficulty alone, attaching an eligible-model policy to the kind of file a subtask is likely to touch:
| Files likely touched | Policy | Eligible models |
|---|---|---|
| Public docs, open-source dependencies | open | any, cheapest first |
| Application code | standard | vetted providers |
| Secrets, environment, infrastructure config | restricted | first-party frontier only |
| Proprietary research code | custom | excludes named vendors |
The underlying point generalizes beyond that specific table: once routing is policy-driven rather than ad hoc, a preference to avoid a given vendor for a class of files becomes configuration the harness enforces automatically, rather than discipline every engineer must remember.
Best practice: permission checks must occur before model routing. Confidence can tighten policy; it must never grant authority the caller does not already possess.
Method 9: Share retrieval, then fan out independent review
Background reviewers are useful when they add different judgments, not when each one repeats the same expensive discovery pass.
A better structure is:
Retrieve and normalize evidence once.
Freeze an evidence-pack version.
Ask multiple independent questions over that same state.
Send only tasks needing generation to LLM reviewers.
Merge structured results by policy.
TypeSafe calls this speculative fan-out: ask independent questions together, including questions that may only matter on one branch, and let code ignore irrelevant answers.
Good fan-out questions for a code patch
Does the patch address the reproduced failure?
Does it modify unrelated behavior?
Does it touch an authorization boundary?
Does it add a dependency?
Does it contradict repository instructions?
Does it need human approval?
Each question should be narrow and independent. TypeSafe explicitly warns that broad questions hide multiple judgments.
Best practice: immutable evidence packs make reviewer disagreements diagnosable. If two reviewers saw different state, their disagreement is not meaningful.
Method 10: Gate every consequential command
A model should not be the final authority over shell commands, deployments, data deletion, or production mutation.
A robust gate has three outcomes:
allow: low-risk and within existing authority
ask: requires human approval or clarification
deny: prohibited regardless of model confidence
Jev can contribute semantic signals:
Does the script delete data?
Does it access secrets?
Does it modify production infrastructure?
Does it match the stated task?
Is a rollback path present?
Code and infrastructure must make the final decision.
assessment = client.system_one(
state={
"task": task,
"command": command,
"working_directory": cwd,
"policy": active_policy,
},
questions={
"matches_task": Noul(
instructions="Does the command directly support the stated task?"
),
"is_destructive": Noul(
instructions="Could the command delete, overwrite, or irreversibly mutate data?"
),
"risk": Score(
instructions="How much damage could an incorrect execution cause?",
criteria=[
"Local and immediately reversible",
"Limited impact with a tested rollback",
"Production, security, financial, or irreversible impact",
],
),
},
)
answers = assessment.answers
if hard_policy_denies(command, identity, cwd):
decision = "deny"
elif answers["is_destructive"].noul > 0.45:
decision = "ask"
elif answers["matches_task"].noul < 0.70:
decision = "ask"
elif answers["risk"].score > 1.25:
decision = "ask"
else:
decision = "allow"
TypeSafe's LLM guardrails cookbook follows the same division: Jev supplies assessments; application policy maps them to pass, review, block, or support. The same probabilities can produce different actions under different named policies.
An independently compiled synthesis of the design notes expresses the deterministic half of that gate as a small policy language, with Jev's semantic signals feeding the conditions rather than replacing them:
policy "exec":
deny if command touches ~/.ssh or .env*
deny if script contents contain network egress
and task.scope != "deploy"
ask if command writes outside repo root
allow if command in read_only_set
allow if tests/ and exit code is expected
The same source recommends inspecting script contents, not only the command name, before execution, for example reading a Python or shell file before running it rather than approving the executable alone. This is a design pattern to adapt, not an official Jev policy DSL.
What a command gate must log
requested command and normalized script
requesting identity
active task and policy version
Jev model version
raw probabilities and scores
deterministic policy checks
final decision
approving human, if any
execution result and side effects
Best practice: inspect the full script and arguments, not only the tool name. terminal is not inherently dangerous; terminal("Remove-Item -Recurse ...") may be.
The complete harness loop
The ten methods combine into a runtime in which state, decisions, execution, and verification remain separate.
Reference pseudocode
while not task.finished:
state = store.load(task.id)
candidates = retrieve(state.goal, state.repo_snapshot)
context_signals = jev_score_context(state.goal, candidates)
context = assemble_context(candidates, context_signals, token_budget)
route = jev_route(state, context, available_handlers)
proposal = coding_llm.run(state.goal, context, route)
for action in proposal.actions:
policy = deterministic_policy(action, identity, environment)
semantic = jev_assess_action(state, action, policy)
decision = combine_policy(policy, semantic)
if decision == "deny":
store.record_denial(task.id, action, policy, semantic)
continue
if decision == "ask":
if not human_approval(action, policy, semantic):
continue
observation = execute_in_sandbox(action)
store.append_observation(task.id, observation)
validation = run_targeted_checks(task)
verification = jev_verify_trace(store.trace(task.id), validation)
task.finished = completion_policy(validation, verification)
This pseudocode is an architectural example, not an official Jev harness SDK.
Failure modes and design limits
TypeSafe publishes a Jev 1.13 jaggedness page. That candor is important because typed output is not the same as semantic correctness.
Documented limitations include:
literal interpretation
unreliable counting and arithmetic
weak date ordering and comparison
reduced accuracy with indirection
context rot from irrelevant state
susceptibility to adversarial content
confusion from contradictory instructions and criteria
no guaranteed arithmetic invariants between separately asked questions
no text generation
Engineering responses
| Limitation | Harness response |
|---|---|
| Math and counting | Compute deterministically in code |
| Dates | Extract components, then compare in code |
| Irrelevant state | Retrieve and filter before calling Jev |
| Adversarial content | Treat state as untrusted; test and enforce policy outside the model |
| Literal reading | Write explicit boundaries and examples |
| Structural inconsistency | Do not assume separate probabilities obey identities |
| Generation | Use an LLM for text and code generation |
| Confidence | Validate thresholds on domain data; do not treat confidence as truth |
Type safety guarantees that the answer matches the declared shape. It does not guarantee that the judgment is correct.
How to evaluate a Jev-assisted coding harness
A production evaluation should compare systems, not isolated model calls.
Quality metrics
relevant context retained
irrelevant context removed
correct skill or tool suggested
false command allows
false command blocks
task completion rate
regression rate
human-review precision and recall
Cost and performance metrics
total tokens across all models
Jev input and output usage
number of model round trips
cache hit rate
end-to-end latency
cost per verified task
Reliability metrics
variance across repeated runs
recovery after tool failure
restart recovery from durable state
policy adherence after context compaction
performance on adversarial retrieved text
calibration of confidence against observed accuracy
An ablation plan
Run the same task set with:
Baseline coding agent.
Agent plus query-aware context scoring.
Agent plus progressive tool disclosure.
Agent plus confidence-aware routing.
Agent plus command gating.
Complete harness.
This shows which components add value and which only add complexity.
What the ten methods solve
| Inherited problem | Engineering response |
|---|---|
| Repeated context processing | Durable state and bounded context assembly |
| Blind model routing | Measured, confidence-aware routing |
| Tool-schema bloat | Progressive disclosure and shortlist reranking |
| Blind compaction | Query-aware chunk judgment |
| Implicit sub-agent state | Versioned evidence packs |
| Lost state on restart | Explicit persisted task records |
| Unsafe execution | Deterministic policy plus semantic command assessment |
| Reviewer duplication | Shared retrieval and parallel independent judgments |
| Unclear uncertainty | Probabilities, confidence, and review paths |
| Opaque workflow | Typed decisions and auditable code-owned control flow |
Current public applications and experiments
Jev's public ecosystem is still early. The examples below are applications that could be verified through TypeSafe's official documentation, cookbooks, evaluation site, launch material, or public repositories as of September 2026. They should not be read as a list of independently confirmed customer production deployments.
Interactive and runnable demos
| Application | What Jev does | Evidence level | Source |
|---|---|---|---|
| Smart-home assistant | Classifies request category, room, device, and action; falls back to an LLM for conversation or compound-request splitting | Official interactive demo | Smart-home assistant |
| Natural-language trading interface | Selects one of ten typed functions and fills bounded arguments from natural language | Reproducible official cookbook | Function calling |
| Jev Playground | Evaluates user-provided state with Choice, Score, and Noul questions | Official hosted application | TypeSafe Playground |
The smart-home demo is a particularly clear example of the intended architecture. Jev handles bounded decisions in parallel; code filters the answers; an LLM is invoked only when the request requires free-form generation.
Agent and developer-tooling applications
| Application | What Jev does | Published result or behavior | Source |
|---|---|---|---|
| Hermes skill suggestion | Ranks 182 skills, reranks a shortlist, and can recommend no skill | In the published 488-request experiment, wrong loads fell from 16.8% to 7.3% and needless loads from 9.8% to 4.0% | Skill suggestion |
| TypeSafe agent skill | Gives Claude Code, Codex, and other coding agents current TypeSafe integration guidance | Public installable skill; it helps agents write TypeSafe integrations but does not replace their LLM | Agent skill repository |
| System One adapter | Runs the same TypeSafe-shaped questions through OpenAI-compatible or Anthropic LLM APIs | Public comparison and compatibility tool with retries and diagnostics | System One adapter |
The Hermes experiment is the closest verified public example to the coding-harness design in this article. Jev does not load the skill or run the agent. It supplies a bounded suggestion that the existing agent may follow or ignore.
Retrieval, verification, and safety applications
| Application | Jev's role | Source |
|---|---|---|
| RAG passage gating | Scores relevance, usable evidence, premise contradiction, and prompt injection before generation | Classifying RAG passages |
| Citation verification | Classifies whether source context supports, contradicts, or says nothing about a claim | Double-checking citations |
| LLM guardrails | Screens inputs and outputs for jailbreaks, harmful requests, medical advice, self-harm, and severity | Guardrails for LLMs |
| Legal retrieval reranking | Reranks BM25 candidates for legal queries | Re-ranking cookbook |
| Knowledge-graph entity alignment | Judges whether candidate records represent the same entity and surfaces disagreeing fields | Entity alignment |
| Structured extraction cascade | Verifies extracted fields and escalates difficult cases to a reasoning model | SDE cascade |
These are worked examples, not generic claims. Their thresholds, model versions, datasets, and reported results belong to those specific experiments.
Published workflow evaluations
TypeSafe's workflow evaluation site publishes four larger policy-shaped applications:
Security incidents: decide whether to close an alert, send it to an analyst, or contain a machine.
Agent trace observability: inspect a completed support-agent trace and decide whether human review is needed and how urgently.
Invoice processing: decide whether to pay, hold, or return a vendor invoice using the invoice, purchase order, and delivery state.
Customer service: decide what a support system should do next from the conversation and account state.
The evaluation assumes the workflow itself is correct and compares model judgments inside that fixed compute graph. It is evidence about structured workflow performance under TypeSafe's published methodology, not independent proof that the workflows are production deployments.
Launch demonstrations
The creator-authored Jev launch article also describes:
a Doom bot making rapid decisions from structured textual game state
Wikipedia racing, where Jev repeatedly chooses links from large candidate sets
These demonstrations illustrate latency, bounded choice, and repeated decision-making. TypeSafe notes important caveats: the Doom demo uses structured text rather than vision, and the Wikipedia-racing comparison settings affect the observed speedups.
What could be verified from X
The official TypeSafe AI X account is public and identifies the company as an AI lab building intelligence beyond chat. However, unauthenticated access during this review exposed the profile but not searchable post bodies. X search pages required login, so this article does not attribute any application, performance result, or third-party deployment to an X post that could not be read and linked directly.
No independently verifiable third-party production deployment was identified in the public sources reviewed here. That does not mean none exist; it means the evidence currently available for this article is strongest for TypeSafe's own runnable demos, cookbooks, SDKs, repositories, and published evaluations.
Conclusion
The most important idea in Jev engineering is not that a smaller model should replace a larger one.
It is that generation and decision-making do not need to be the same operation.
A generative coding model is useful because software work is open-ended. It must read unfamiliar code, construct hypotheses, explain failures, and create patches. But many decisions inside that workflow are bounded:
choose one tool
score one risk dimension
judge whether one condition holds
rank candidate context
decide whether uncertainty requires escalation
Jev is designed for that bounded layer.
The architecture becomes easier to reason about when:
state is explicit
evidence has provenance
questions are atomic
probabilities remain visible
code owns thresholds and control flow
infrastructure owns authorization
the LLM owns generation
humans own consequential approvals
That is the deeper lesson behind the Jev harness idea:
Do not ask one model to remember, decide, generate, authorize, execute, and verify everything. Give each responsibility to the component that can make it explicit, testable, and governable.
Jev's opportunity is therefore architectural. It offers a machine-oriented decision interface at points where agent systems often rely on hidden prompt logic: context selection, routing, tool choice, verification, risk assessment, and escalation. Its typed outputs make those decisions easier to inspect and compose, while probabilities make uncertainty available to policy code.
Its limits are equally important. Jev does not make authorization unnecessary, does not turn model confidence into truth, does not remove prompt-injection risk, and does not replace a generative coding model. Its current documented jagged edges include literal reading, context rot, numeric weakness, date-comparison weakness, adversarial influence, and lack of text generation.
The strongest implementation path is incremental:
Pick one bounded decision already causing cost, latency, or reliability problems.
Define the state and one atomic typed question.
Keep deterministic rules and side effects in code.
Validate probabilities and thresholds on representative data.
Add a human or reasoning-model route for uncertain cases.
Measure the complete workflow, not only the Jev call.
If that first decision improves the system, add another. The goal is not to place Jev everywhere. The goal is to replace opaque, repeated, free-form judgment with small decision interfaces wherever doing so makes the harness more reliable, economical, and governable.
