How NOX Tests Leo
Leo is tested as a safety layer, not treated as a promise that software can catch every emergency. Nox measures the deterministic red-flag detector against a maintained test set and publishes high-level recall and false-positive-rate results on its Trust & Transparency page.
That distinction matters. A health AI system should be open about what is measured, what is not included in those measurements, and why a safety warning must appear before a conversational answer when an acute red flag is detected.
What does Leo test for?
Leo screens messages for signs of acute red-flag conditions before the AI model answers. Its deterministic triage layer includes more than 100 rules across more than 70 acute red-flag categories, including stroke signs, chest pain and cardiac symptoms, severe breathing difficulty, severe allergic reactions, trauma and severe bleeding, mental-health crisis, pregnancy warning signs, concerning symptoms in children, environmental emergencies, and toxicology or overdose.
When a deterministic rule fires, Nox displays a clear emergency or urgent-care banner before AI-generated content. The purpose is not to decide what condition a person has. It is to recognize descriptions that may require prompt professional or emergency care and make that guidance visible without relying on a generated response to choose the right moment or wording.
This is why safety testing cannot be reduced to whether an AI answer sounds knowledgeable. A response can be articulate, empathetic, and factually plausible while still failing to foreground urgent next steps when a message describes a potential emergency. For more on that distinction, read Why High Model Accuracy Is Not Enough for Health AI.
How Nox measures Leo’s deterministic detector
Nox measures the deterministic detector’s overall recall and false-positive rate against a maintained test set. Those high-level summary metrics are published publicly on the Trust & Transparency page, with machine-readable metrics also available through Nox’s transparency materials.
Recall asks an essential safety question: among red-flag cases in the test set, how often did the detector identify them? A false-positive rate addresses a different practical question: how often does the detector raise a warning when the tested message does not meet its red-flag criteria?
Neither measure tells the whole story by itself. A system that rarely warns may avoid unnecessary alerts but miss messages that need attention. A system that warns very broadly may catch more concerning cases but can also interrupt ordinary conversations more often. Publishing both measures makes the tradeoff more visible than a single headline score.
Importantly, the published metrics apply to the deterministic detector. They do not represent a claim that every part of Leo has the same measured performance, nor do they mean Leo catches every emergency. Nox explicitly describes Leo as a safety net, not a guarantee.
Why deterministic rules are tested separately
Leo’s first screen is independent and deterministic: it uses fixed rules that run before any AI model is called. That design makes the layer measurable in a direct, repeatable way. Given the same message and settings, the deterministic layer follows the same rules.
This is not an argument that fixed rules can understand every way a person might describe a dangerous situation. People use shorthand, slang, indirect language, older disease names, and many languages. A phrase can also become concerning only in the context of recent messages.
But deterministic screening gives Nox a safety mechanism whose recall and false-positive rate can be evaluated and published separately. It also ensures that a detected warning is not dependent on a conversational model deciding whether to mention emergency care in its reply. Learn more about the design in How Leo Works.
What happens beyond the deterministic test
If the deterministic layer finds no match, Leo can run a second check in the background: a lightweight AI classifier re-reads recent messages for dangerous descriptions that fixed patterns can miss. This backstop is intended to recognize likely emergencies expressed indirectly, in slang, in another language, or through older terminology.
If the classifier recognizes a likely emergency, Nox surfaces the same seek-care guidance used for a deterministic pattern match. Crucially, this backstop can add a safety note only; it cannot remove or soften one that has already been triggered.
Because that second check is not deterministic, its results are kept separate from the published deterministic-detector metrics. That separation is important for honest reporting. Combining different systems into one number can make it hard to tell what was tested, how it was tested, and what the number actually represents.
How Leo’s settings affect screening
Leo gives people a choice of screening strictness directly in the chat box. Relaxed warns about clear emergencies only. Standard also warns about anything urgent. Strict adds the AI backstop that reviews recent messages for dangerous descriptions that are indirect or expressed in another language.
Whatever setting a person chooses, Leo continues to screen for true emergencies. The setting changes how many optional layers run on top of that core emergency screening, rather than allowing emergency detection to be switched off.
Agent and voice conversations always use the strictest setting. This reflects a safety-first approach in contexts where a person may be speaking naturally, moving quickly, or describing a situation less formally than they would in typed text.
What a safety warning should do
When Leo recognizes a likely emergency, its job is to surface clear guidance to seek appropriate care. If a region is set in Settings, the banner shows the local emergency number. If no region is available, Nox falls back to universal emergency numbers—911, 999, or 112—so that the guidance is not blank.
The warning arrives before any AI-generated content. That order is deliberate: in a potentially urgent situation, the reader should not have to scan through an explanation to find the most important direction.
If you or someone else may be experiencing an emergency, contact local emergency services. For symptoms that are concerning, worsening, or persistent, a qualified clinician can provide an assessment that an educational AI companion cannot.
Why published metrics are only one part of accountability
Testing and publishing metrics are meaningful steps, but they do not transform an AI companion into a medical device or a substitute for clinical judgment. Test sets are limited representations of real-world language and real-world risk. They cannot prove that every future message, phrasing, language variation, or situation will be handled correctly.
That is why the strongest safety approach combines measurement with clear boundaries. Nox does not claim Leo catches every emergency. It identifies the deterministic layer that is measured, keeps non-deterministic backstop results separate, and puts emergency guidance ahead of generated answers when a likely red flag is detected.
For readers evaluating any health AI, the useful questions are straightforward: Does safety screening happen before the response? Is there a clear path for urgent escalation? Are performance measures public? And does the product explain the limits of those measures? Explore a broader framework in How to Evaluate the Safety of an AI Health Assistant.
Common questions
Does Nox test every part of Leo with the same published metric?
No. Nox publishes overall recall and false-positive-rate summaries for the deterministic red-flag detector against a maintained test set. The AI classifier backstop is separate from those published deterministic metrics.
Does a high recall score mean Leo catches every emergency?
No. Recall describes performance on the maintained test set. Leo is a safety net, not a guarantee, and Nox does not claim that it catches every emergency.
Why does Leo show guidance before an AI answer?
A red-flag warning is designed to make urgent or emergency guidance visible first, rather than relying on a generated response to mention it later.
Can Leo recognize indirect or multilingual descriptions?
The deterministic layer includes multilingual coverage. On Strict, Leo also uses an AI classifier backstop that can review recent messages for likely emergencies expressed indirectly, in slang, or in another language.