Verifiable AI work.

AI is starting to do real work. That work should leave evidence and be checkable by something other than the system that did it. That idea organizes what Nova Labs builds.

“Done” is not evidence.

An AI system can report that it finished a task and still have failed the last step: the file was not saved, the change did not persist after a refresh, the form went out with the wrong value.

When AI only answered questions, a person read the answer and judged it. When it acts in applications, the browser and files, the result is no longer in the text it returns. It is in the state of the system after it acted, and that is where you have to look.

Execute. Leave evidence. Verify.

Three stages, one responsibility each. Our products sit inside them.

Dovren executes. Tesven verifies: one does the work, the other checks with evidence whether it was done right.

How the work is split across Nova Labs products · status as of October 2026
  1. ExecutionAn AI system receives a goal and acts with real tools: applications, the browser, files.Dovren · In development
  2. EvidenceEach action leaves a reviewable record: what was done, under which permission, and what result remained.Dovren and Tesven · Activity log and evidence bundles
  3. VerificationSomething other than the executor checks the result against the criteria defined for the task and returns a verdict, including “could not be verified”.Tesven · Release candidate v0.1.0

The executor does not grade itself.

If the system that did the work also decides whether it was done right, its mistake and its check share a cause: if it misread the task, it will check it the same wrong way.

So Tesven does not ask the executor whether it finished. It reads the resulting state on its own, through independent reads of authorized evidence, and compares it with the expected outcome defined for the task. It is the same principle that separates the person who writes a program from the person who tests it, or the clerk who records a transaction from the auditor.

A verdict that can say no.

Tesven returns one of five verdicts. “Unverified” matters as much as “Verified”: missing evidence is never treated as success, and a verified result only holds for the checks that ran.

Verified
VERIFIED_PASS
Failed
VERIFIED_FAIL
Unverified
INSUFFICIENT_EVIDENCE
Human review
HUMAN_REVIEW
Not applicable
NOT_APPLICABLE
How Tesven verifies

What exists today, and what does not.

Dovren
In development. No verified public download.
Tesven
Release candidate v0.1.0 for local technical evaluation with synthetic data. No public consumer installer.
Pairing
Today, Dovren evidence can be imported into Tesven manually; automatic pairing is planned.
Success rates
We do not publish success rates for Dovren or Tesven. There is no public, versioned task set behind them yet. When there is, it will ship with methodology, environment, criteria and limits, as PY-AIEval did.

The same idea across Nova.

Nova Research

PY-AIEval keeps the configuration, outputs, scores and hashes of every run so anyone can re-score it offline.

PY-AIEval in English

Nova Standard v0.1

Eight controls before widening an agent’s autonomy, including observability, verification and exception handling. In Spanish.

Read the standard

Systems for companies

Every delivery is defined with acceptance criteria and checked before it goes into use. In Spanish.

How we work

Questions we are still studying.

  1. How do you define acceptance criteria for an ambiguous task without making the check as expensive as the task?
  2. Which part of the final state can be read independently when an application offers no API?
  3. When should a verifier escalate to human review instead of deciding?

We do not have settled answers. We will publish them as technical notes, with their limits, when we have evidence.

Want to evaluate Tesven or Dovren?

Tell us what work you want to delegate or check. We will tell you what can be tested today and with which limits.

Ask about the products