Nox 1.4 Orion Preview (Preview)
The newest model — longer, more complete answers, in preview
Today, we're introducing 1.4 Orion Preview — the first release of the Nox 1.4 generation, offered as a preview alongside the 1.3 family. Orion is the first Nox model whose answer length isn't fixed: it sizes each reply to the question, from a few sentences to a full 9,000-character answer, reasoning through a deliberate internal pass with a roughly one-million-token context window before it responds.
The headline change is length. Every earlier Nox model answered under a fixed 4,000-character cap. In a paired experiment on 150 HealthBench Hard examples graded with the official grader, lifting that cap alone raised the score by 8.4 ± 2.0 points, almost entirely on the completeness axis, without a drop in accuracy. Orion scales its length to the question instead — short for simple questions, up to 9,000 characters when a question genuinely has several parts — and is explicitly asked to cover the direct answer, the main alternatives, next steps, and what would warrant urgent care.
Two further parts of the answer standard target where the 1.3 family scored lowest: answers addressed to clinicians give protocol-level specifics rather than lay-level hedging, and possibly emergent situations get concrete first actions — who to call, what to do while waiting, what not to do — before anything else.
On frontier reasoning and science tests, the engine behind Orion is reported to answer a meaningful share of Humanity's Last Exam correctly and to score in the high 80s on a graduate-level science benchmark spanning biology, chemistry, and physics — see the chart below for the figures and how they sit against the wider field.
On HealthBench Hard — the 1,000-prompt hard subset of the open, physician-built HealthBench — 1.4 Orion Preview scored 46.4% (±0.8) in Nox's own self-run on September 5, 2026: all 1,000 examples and all 11,846 clinician-written rubric criteria, graded with GPT-4.1, the original official HealthBench grader model, on the official open protocol. The system under test was the full Nox production pipeline with Leo active, at Orion's strongest user-selectable settings. On the identical protocol and grader, 1.3 Astra scored 34.8%, so Orion is 11.6 points ahead on the same scale. By rubric axis: communication quality 60.7%, accuracy 59.1%, instruction following 55.6%, completeness 43.5%, context awareness 32.8%. By theme: emergency referrals 55.9%, hedging 50.8%, health data tasks 48.9%, context seeking 46.0%, global health 45.3%, communication 42.0%, complex responses 36.5%. Mean answer length was 3,869 characters; emergency guidance appeared on 19.4% of examples; no answer was truncated. The complete raw record — all 1,000 original answers, all 11,846 grader verdicts with written reasoning, the aggregate summary, and SHA-256 checksums — is published at https://www.getnox.ai/data/healthbench/orion/ (README.md there documents the schemas and how to recompute the score), and the Orion page walks through the grading process with three complete graded examples. Self-run scores are disclosed in full but not independently verified, and a HealthBench score measures performance against that benchmark's rubric, not overall product quality or clinical outcomes.
Orion is multimodal within Nox and works with tools. It takes photo input — up to ten images can be attached to a single message and are read automatically in the same conversation. It has live web access: in Plan mode Orion searches the web and opens pages on its own, up to ten rounds of searching and reading per answer, and every source it draws on is cited inline. It can generate images — health and anatomy illustrations produced on request, rendered directly in the answer. And its reasoning is visible: the time Orion spends thinking is shown on every answer, and Plan mode shows a live step-by-step progress timeline while it works.
In NOX WORK, the same Orion model additionally reads attached documents — PDF, Word, Excel, PowerPoint, CSV, plain text, Markdown, JSON, and code, up to three files of 15 MB each per message — returns tabular results as tables, can remove backgrounds from and upscale uploaded images, and reads from connected apps such as Google Calendar, Drive, and Todoist. Live web is on by default there.
Like every model in the family, Orion runs behind Leo, Nox's deterministic safety layer, on every message. Web access, image generation, and app actions are all withheld on any turn Leo flags as a possible emergency, so the emergency guidance is never delayed by a tool call.
1.4 Orion Preview is available in Plan mode on Pro and MAX, and as a selectable model in NOX WORK. It is not an everyday-chat model; 1.3 Astra remains the flagship while Orion is in preview.
Capabilities
- Deliberate multi-step reasoning — an internal chain of thought worked through before every answer, with thinking time shown
- ~1,000,000-token context window
- Dynamic answer length — sized to the question, up to a 9,000-character hard cap
- Full Nox answer standard: completeness check, audience matching, concrete emergency actions
- Photo input — up to ten images per message, read automatically
- Live web access in Plan mode — searches and opens pages itself, up to ten rounds per answer, every source cited
- Image generation — health and anatomy illustrations rendered inline on request
- Plan mode progress timeline — each research and reasoning step shown live as it happens
- In NOX WORK: document reading (PDF, Word, Excel, PowerPoint, CSV, text, code), image background removal and upscaling, connected-app reads
Model details
Availability: Available in Plan mode on Pro and MAX.