Verifiable AI work.
AI is starting to do real work. That work should leave evidence and be checkable by something other than the system that did it. That idea organizes what Nova Labs builds.
“Done” is not evidence.
An AI system can report that it finished a task and still have failed the last step: the file was not saved, the change did not persist after a refresh, the form went out with the wrong value.
When AI only answered questions, a person read the answer and judged it. When it acts in applications, the browser and files, the result is no longer in the text it returns. It is in the state of the system after it acted, and that is where you have to look.
Execute. Leave evidence. Verify.
Three stages, one responsibility each. Our products sit inside them.
Dovren executes. Tesven verifies: one does the work, the other checks with evidence whether it was done right.
- ExecutionAn AI system receives a goal and acts with real tools: applications, the browser, files.Dovren · In development
- EvidenceEach action leaves a reviewable record: what was done, under which permission, and what result remained.Dovren and Tesven · Activity log and evidence bundles
- VerificationSomething other than the executor checks the result against the criteria defined for the task and returns a verdict, including “could not be verified”.Tesven · Release candidate v0.1.0
The executor does not grade itself.
If the system that did the work also decides whether it was done right, its mistake and its check share a cause: if it misread the task, it will check it the same wrong way.
So Tesven does not ask the executor whether it finished. It reads the resulting state on its own, through independent reads of authorized evidence, and compares it with the expected outcome defined for the task. It is the same principle that separates the person who writes a program from the person who tests it, or the clerk who records a transaction from the auditor.
A verdict that can say no.
Tesven returns one of five verdicts. “Unverified” matters as much as “Verified”: missing evidence is never treated as success, and a verified result only holds for the checks that ran.
- Verified
VERIFIED_PASS- Failed
VERIFIED_FAIL- Unverified
INSUFFICIENT_EVIDENCE- Human review
HUMAN_REVIEW- Not applicable
NOT_APPLICABLE
What exists today, and what does not.
- Dovren
- In development. No verified public download.
- Tesven
- Release candidate v0.1.0 for local technical evaluation with synthetic data. No public consumer installer.
- Pairing
- Today, Dovren evidence can be imported into Tesven manually; automatic pairing is planned.
- Success rates
- We do not publish success rates for Dovren or Tesven. There is no public, versioned task set behind them yet. When there is, it will ship with methodology, environment, criteria and limits, as PY-AIEval did.
The same idea across Nova.
Nova Research
PY-AIEval keeps the configuration, outputs, scores and hashes of every run so anyone can re-score it offline.
PY-AIEval in EnglishNova Standard v0.1
Eight controls before widening an agent’s autonomy, including observability, verification and exception handling. In Spanish.
Read the standardSystems for companies
Every delivery is defined with acceptance criteria and checked before it goes into use. In Spanish.
How we workQuestions we are still studying.
- How do you define acceptance criteria for an ambiguous task without making the check as expensive as the task?
- Which part of the final state can be read independently when an application offers no API?
- When should a verifier escalate to human review instead of deciding?
We do not have settled answers. We will publish them as technical notes, with their limits, when we have evidence.
Want to evaluate Tesven or Dovren?
Tell us what work you want to delegate or check. We will tell you what can be tested today and with which limits.
