PY-AIEval v0.1.0, in short.
An open methodology and harness from Nova Labs Research for reproducible AI evaluation with a Paraguayan focus. Development stage: no leaderboard, no model-superiority claims.
What it is.
PY-AIEval separates task definitions, provider execution, stored outputs, deterministic scoring and publication gates. Each run keeps its configuration, outputs, scores and SHA-256 hashes, so it can be audited and re-scored offline.
Version 0.1.0 was released as open source (MIT) on 3 October 2026.
What is included.
- A task schema and deterministic scorers.
- A provider-neutral runner with two adapters under one contract: OpenAI Responses and Hugging Face Chat.
- Comparability, contamination and release gates.
- Four development task sets: reasoning, factual knowledge about Paraguay, Paraguayan Spanish, and structured tool use. All are
rankable: false. - One reproducible system-baseline example (Stage 1).
Stage 1: was model training needed?
A deterministic retrieval baseline scored 5/5 on a factual slice of five items, with 0 errors, 0 external LLM or API calls and 0 training compute. The success criterion was fixed before the final run. The first implementation scored 4/5 because of a retrieval-ranking collision; the fix and its regression test are documented in the corrections log (in Spanish).
The result supports one narrow engineering decision: not to train or fine-tune a model for that factual slice.
How to cite.
@software{novalabs_pyaieval_2026,
author = {{Nova Labs Research}},
title = {PY-AIEval},
version = {0.1.0},
year = {2026},
month = oct,
url = {https://www.novalabs.com.py/investigacion/benchmark},
repository = {https://github.com/Nova-Labs-Paraguay/novalabs-py-aieval},
license = {MIT},
note = {Development task sets (rankable: false); no DOI}
}The same metadata is in the repository’s CITATION.cff. There is no DOI and no peer review.
What it does not show.
- It is not a leaderboard and does not compare models.
- It does not show general knowledge of Paraguay, nor Guaraní or Jopara competence.
- It does not claim to be the first, best or only Paraguayan benchmark.
- Nova did not train a model of its own for this release.
- The task sets are small and visible, so they are unsuitable for ranking claims.
Verify, critique and cite.
Clone the repository and run npm install, npm run typecheck, npm test and npm run verify:release. To replicate or criticize a result, open an issue with the release version, protocol and evidence.
Cite as: Nova Labs Research. PY-AIEval v0.1.0. 2026. https://www.novalabs.com.py/investigacion/benchmark. There is no DOI for this version.
Questions about PY-AIEval?
Write to [email protected]. We answer in English or Spanish.
Diagnóstico gratuito: 6 preguntas y una recomendación preliminar. Al final te pedimos nombre y correo.
