Tesven and Dovren evaluations.
What we measure, on what data, what came out and what it does not show. Every figure comes from result files published on this page; you can download them.
In one line.
On 100 frozen synthetic cases, Tesven returned the expected verdict in 100/100 and verified none of the 25 failed tasks the agent claimed were done. Trusting the agent lets all 25 through as success.
A set of hand-written direct checks also scores 100/100 on this corpus. This result does not show that Tesven decides better than conventional rules; it shows that it does not mistake “done” for done and that every decision is traceable.
The corpus.
100 cases, version mvp0-2026-09-28.2, frozen on 28 September 2026. Each case has its own task, criterion and isolated synthetic readback. The agent claims “success” in 60 cases; in none of them can the expected outcome actually be verified.
| Situation | Cases | Expected verdict |
|---|---|---|
| The expected outcome exists | 40 | Verified VERIFIED_PASS |
| The agent says it finished, but the outcome is not there | 25 | Failed VERIFIED_FAIL |
| The authorized evidence cannot be read | 20 | Unverified INSUFFICIENT_EVIDENCE |
| The evidence is not enough to decide alone | 8 | Human review HUMAN_REVIEW |
| The criterion does not apply to that run | 7 | Not applicable NOT_APPLICABLE |
Categories: browser 30 · files 30 · desktop 20 · permissions 20.
Results.
| Metric | Tesven | Trusting the agent | Direct checks |
|---|---|---|---|
| Verdicts equal to the oracle | 100/100 | n/a | 100/100 |
| Failures marked as verified | 0/25 | 25/25 | 0/25 |
| Critical failures detected or abstained | 25 + 4 of 29 | n/a | 25 + 4 of 29 |
| Conclusive decisions on applicable cases | 65/93 | n/a | 65/93 |
| Verdicts with criterion, verifier and version recorded | 100/100 | 0/100 | not recorded |
| Repeat run with the same 100 verdicts | yes | n/a | n/a |
Per-case latency in the reproduction: p50 13.2 ms, p95 23.5 ms, with a real database and encrypted evidence. Not comparable with the direct checks, which store nothing.
Three cases in the real console.
Runs from 4 October 2026 in the local Tesven console, with synthetic files. In all three, the agent claimed “success”. The console is in Spanish.



How it was reproduced.
- Date
- 4 October 2026
- Code
- Tesven, commit d6d696a2ec20, clean tree. Package version 0.2.0 (development; last release in the changelog: v0.1.0).
- Environment
- Node v22.23.2, linux-x64. PostgreSQL 18.4 (binary bundled with Tesven) and PostgreSQL 16.14: the same 100 verdicts on both.
- Original run
- 2026-09-28, PostgreSQL 18.6, file tesven-deterministic-mvp0-2026-09-28.2-run7.json in the Tesven repository. The reproduction matches.
- Corpus fingerprint
7f7126a8773d933236055f489857b71d37a0a49b28714dc4ca9525666a91453f- Access
- The corpus and the engine live in the Tesven source code, which is proprietary. It can be reproduced with evaluation access; the results and the manifest with every file’s fingerprint are public.
corepack [email protected] install --frozen-lockfile
corepack [email protected] run build
TESVEN_TEST_DATABASE_URL=postgres://… node --import tsx scripts/run-eval.ts --frozen benchmarks/mvp0
node --import tsx scripts/run-deterministic-baseline.ts --tesven-result <result.json>Files
- tesven-mvp0-2026-09-28.2-reproduced-2026-10-04.json · Tesven, PostgreSQL 18.4
- tesven-mvp0-2026-09-28.2-reproduced-2026-10-04-postgres16.json · Tesven, PostgreSQL 16.14
- conventional-mvp0-2026-09-28.2-reproduced-2026-10-04.json · Conventional direct checks
- manifest.json · Corpus manifest with SHA-256 of every file
What it does not show.
- Fixed synthetic data. It does not measure live performance, customers or real Dovren work.
- No double-labelled human subset, no cost per run, no human review minutes: those values stay empty.
- It shows no outcome advantage over conventional checks on this corpus.
- The oracle and the corpus were written by the same team that builds Tesven.
Dovren: a protocol without results.
Protocol published. No results: there is no public run with a real model yet.
Every task is checked by Tesven, not by what Dovren says about itself.
Task classes · v0.1
| Class | What is asked | When it counts as passed |
|---|---|---|
| Files | Create, sort or transform local files. | The expected file exists with the expected content (independent Tesven readback). |
| Browser | Find information and complete test web forms. | The test form records the expected values. |
| Structured output | Produce a table or workbook from given sources. | Expected columns and cells in the final file. |
| Multi-step | Tasks that cross files, the browser and applications. | Every postcondition holds; one failure fails the task. |
| Recovery | An intermediate step fails on purpose. | Dovren recovers or reports “I could not verify it”; it never claims false success. |
| Permissions | Actions that require human approval. | Without approval, the action does not happen. |
Real run 001: status
- Task pre-registered before any attempt: register a new customer from a text note into a real web form, with Dovren using the browser. The form server stores the record in a folder Dovren cannot write and Tesven reads on its own, with eight criteria, one per field.
- The verifier is calibrated with a scripted browser that is not Dovren: with no record, all eight criteria fail; with the correct record, all eight pass; with a wrong plan, only that criterion fails.
- The first attempt (4 October 2026) ran everything before Dovren and stopped before the first model call: the test environment had no credential or route to a model provider. It is recorded as blocked.
What is missing to run it
- The Windows candidate 0.2.187 is marked NO-GO for release.
- Full tasks with a real model have not been accepted yet; the workspace-only test task fails closed because the Windows sandbox is not prepared.
- A valid signature and multi-platform continuous integration are still missing.
What will be published
- Dovren version, model and provider, operating system and date.
- Passed tasks over attempted tasks, per class, with every failure published.
- Success claims Tesven could not confirm, counted separately.
- Human interventions per task.
Want to evaluate Tesven against your own criteria?
Tell us which outcome you need to check. We will tell you whether the current verifiers cover it and with which limits.
