Skip to main content

Auryth research lab

Research at Auryth

We publish our methods so you can check for yourself why Auryth gives better answers than generic AI. Others in the field can then build on what we learn.

Our mission

Regulated fields need precision about which rule applied on which date and in which jurisdiction. Generic AI models are not built for that.

That is why general-purpose AI keeps failing on specialist questions.

We publish our research because legal AI gets better when its problems are studied in the open. Hallucination rates, confidence calibration, source attribution and multilingual retrieval are hard problems, and they deserve serious academic attention.

Retrieval research

On their own, large language models are unreliable for high-stakes professional work. Our research focuses on retrieval: the systems that find and check the evidence before the model sees it. Auryth's products are built on patent-pending technology across five areas of retrieval.

Negative evidence in retrieval

Detecting when the evidence contradicts a conclusion or fails to support it.

Calibrated scoring

Confidence scores that match how often answers turn out to be right.

Confidence-gated generation

Output controls that prevent low-confidence answers from reaching users.

Adaptive query routing

Choosing a retrieval strategy based on the kind of question asked.

Self-improving retrieval systems

Feedback loops that improve accuracy without retraining the model.

Research focus areas

Confidence calibration

A confidence score you can rely on

A confidence score is only useful if it matches how often the answer turns out right. We study how to calibrate it so that it does.

Multilingual legal retrieval

Asking in Dutch, finding the answer in French

In a multilingual legal system, the answer is often written in a different language from the question. We study how to find the right provision whatever language it is in.

Temporal versioning

Getting the right rule for the right date

Regulations change all the time. We work on methods to track which version of a provision was in force on the date that matters, so that you don't cite a rule that no longer applies.

Hallucination detection

Catching fabricated citations before they reach you

AI models sometimes cite articles that do not exist, and they do it confidently. We develop methods to check every citation against the real sources before you see the answer.

Working papers

Working paper · Donald Murre

What doesn't match matters more: CRANE, Calibrated Retrieval with Adversarial Negative Evidence

Argues that retrieval for high-stakes domains must represent what a document does not answer, and describes CRANE, an architecture built on typed negative evidence, calibrated confidence and confidence-gated generation. It states four falsifiable predictions for empirical testing.

Download paper (PDF)

Advisory board

We are forming an advisory board of practitioners, academics and AI researchers who care about specialist AI that is transparent and reliable.

We would like to hear from researchers in domain-specific NLP, from academics who study AI in regulated fields, and from practitioners who want to help shape specialist AI tools.

Partnerships

We are looking for partners in three areas:

University research centres

Joint projects on legal AI, NLP, and computational law

Professional associations

ITAA, IBR/IRE, and regional accounting bodies

EU research programmes

Digital governance and AI innovation grants

Work with us

If you would like to work with us on legal AI research, send us a message.

Get in touch