Cernira

For AI labs

Evaluation sets and post-training data written by vetted French specialists in medicine and finance. Bilingual FR/EN, quality-controlled at submission, delivered in your format.

What we produce

Evaluation sets

Experts write the problems: a question, a single verifiable answer, the reasoning steps, a grading rubric, and every source file needed to solve it.

Model evaluation

Accuracy, safety, and completeness scoring of your model outputs against rubrics your team defines.

Preference & RLHF data

Side-by-side comparisons with structured rationales, built for post-training pipelines.

Domain labeling

NER, classification, and structured extraction on clinical and financial text, in French and English.

Who writes it

Medicine

Medical students and junior doctors from French university hospitals.

Finance

Students and analysts in corporate finance, M&A, credit, and IFRS reporting.

Academic email verified, domain declared, and quality scored on every task.

How we keep it usable

Rules enforced at submission

Your format constraints are checked while the expert is still on the page. Work that would be unusable never enters the dataset.

Agreement measured, not asserted

Overlapping annotations produce an inter-annotator agreement score, alongside control tasks with known answers.

Every item traceable

Annotations are immutable and audited: who produced what, when, and against which version of the instructions.

Questions we get asked

How do you know these people are who they say they are?
Every contributor verifies an academic or hospital email address, and the domain is matched against a registry of French universities, teaching hospitals and schools. A recognised domain validates the account; anything else is reviewed by hand. The address proves the institution, so the field a contributor declares is reviewed rather than assumed.
What stops someone submitting work a model could have written?
For authored problems, the brief states a difficulty threshold, typically at most 3 successes out of 16 attempts against a named model, and the author reports the result. A problem a frontier model solves easily has no value in an evaluation set, and the submission form says so while the author is still writing.
How is quality measured rather than claimed?
Tasks can be assigned to several contributors at once, which produces an inter-annotator agreement score. Control tasks with known answers are interleaved with ordinary ones, and every contributor carries a score derived from them. Both numbers are available for the batch you commission.
What happens to the data we send you?
A dataset arrives through a single-use link that expires, and it is validated on arrival: malformed rows are refused with the offending line named, so nothing partial enters the system. Contributors only ever see the items assigned to them, and annotations are immutable and audited.
What if the output is not usable?
Format constraints are enforced at submission, so unusable work does not reach you in the first place. Beyond that, tell us: a batch that misses the brief is a brief we wrote badly, and it gets rewritten rather than invoiced twice.
Tell us about your project
We reply to every request within 24 hours.