Auryth research lab
Research at Auryth
We publish our methods so you can check for yourself why Auryth gives better answers than generic AI. Others in the field can then build on what we learn.
Our mission
Regulated fields need precision about which rule applied on which date and in which jurisdiction. Generic AI models are not built for that.
That is why general-purpose AI keeps failing on specialist questions.
We publish our research because legal AI gets better when its problems are studied in the open. Hallucination rates, confidence calibration, source attribution and multilingual retrieval are hard problems, and they deserve serious academic attention.
Retrieval research
On their own, large language models are unreliable for high-stakes professional work. Our research focuses on retrieval: the systems that find and check the evidence before the model sees it. Auryth's products are built on patent-pending technology across five areas of retrieval.
Negative evidence in retrieval
Detecting when the evidence contradicts a conclusion or fails to support it.
Calibrated scoring
Confidence scores that match how often answers turn out to be right.
Confidence-gated generation
Output controls that prevent low-confidence answers from reaching users.
Adaptive query routing
Choosing a retrieval strategy based on the kind of question asked.
Self-improving retrieval systems
Feedback loops that improve accuracy without retraining the model.
Research focus areas
Confidence calibration
A confidence score you can rely on
A confidence score is only useful if it matches how often the answer turns out right. We study how to calibrate it so that it does.
Multilingual legal retrieval
Asking in Dutch, finding the answer in French
In a multilingual legal system, the answer is often written in a different language from the question. We study how to find the right provision whatever language it is in.
Temporal versioning
Getting the right rule for the right date
Regulations change all the time. We work on methods to track which version of a provision was in force on the date that matters, so that you don't cite a rule that no longer applies.
Hallucination detection
Catching fabricated citations before they reach you
AI models sometimes cite articles that do not exist, and they do it confidently. We develop methods to check every citation against the real sources before you see the answer.
Working papers
Working paper · Donald Murre
What doesn't match matters more: CRANE, Calibrated Retrieval with Adversarial Negative Evidence
Argues that retrieval for high-stakes domains must represent what a document does not answer, and describes CRANE, an architecture built on typed negative evidence, calibrated confidence and confidence-gated generation. It states four falsifiable predictions for empirical testing.
Download paper (PDF)Advisory board
We are forming an advisory board of practitioners, academics and AI researchers who care about specialist AI that is transparent and reliable.
We would like to hear from researchers in domain-specific NLP, from academics who study AI in regulated fields, and from practitioners who want to help shape specialist AI tools.
Partnerships
We are looking for partners in three areas:
University research centres
Joint projects on legal AI, NLP, and computational law
Professional associations
ITAA, IBR/IRE, and regional accounting bodies
EU research programmes
Digital governance and AI innovation grants
Work with us
If you would like to work with us on legal AI research, send us a message.
Get in touch