Nox Orion 2 (Flagship)

The upgrade to 1.4 Orion — faster, more accurate, lower cost

Orion 2 is the direct upgrade to 1.4 Orion and the flagship of the Nox 2 generation. It keeps everything 1.4 Orion could do — the same tools, thinking levels, and answer standard — while responding faster overall, answering more accurately, and costing less. Under the hood it is an 824-billion-parameter mixture-of-experts model, 46 billion active per token, with a roughly one-million-token context window. It reasons through a deliberate internal pass before answering and sizes each answer to the question, up to 9,000 characters.

Architecture: mixture-of-experts, 824 billion parameters total, 46 billion active per token — the 700-billion-parameter backbone of 1.3 Astra extended with roughly 100 billion parameters of additional trainable components.

Length: earlier Nox models answered under a fixed 4,000-character cap. Lifting that cap alone raised HealthBench Hard by 8.4 ± 2.0 points on a 150-example paired test, almost entirely on completeness. Orion scales length to the question instead — short for simple questions, up to 9,000 characters for multi-part ones — and covers the direct answer, main alternatives, next steps, and what would warrant urgent care.

Measured on the 1.4 Orion pipeline in September 2026: HealthBench Hard, Nox self-run, official open protocol, all 1,000 examples, scored 46.4% (±0.8) at Ultra and 53.3% (±1.0) at Hyper Ultra. Orion 2 runs the same engine and answer pipeline, so these remain the most recent published results; they have not been re-measured under the Orion 2 name. 1.3 Astra on the identical protocol: 34.8%.

Hyper Ultra is Orion’s sixth and deepest thinking level, in Plan mode on MAX. Same model, several passes over one question — draft, self-written rubric, self-grade, revision — about five to six minutes per answer. The 1.4 Orion pipeline scored Light 36.2%, Balanced 36.8%, Extra 36.9%, and Max 38.5% on a fixed 150-prompt development subset.

Multimodal and tool-using: up to ten photos per message; live web access in Plan mode with up to ten search-and-read rounds per answer, every source cited; inline health and anatomy image generation; thinking time shown on every answer and a live step timeline in Plan mode. In NOX WORK it also reads attached documents (PDF, Word, Excel, PowerPoint, CSV, text, code — three files of 15 MB each), removes backgrounds from and upscales images, and reads connected apps.

Runs behind Leo, Nox’s deterministic safety layer, on every message; web, image, and app actions are withheld on any turn Leo flags as a possible emergency.

Available in Plan mode on Pro and MAX (Hyper Ultra on MAX) and as a selectable model in NOX WORK. Not an everyday-chat model.

Capabilities

  • 824-billion-parameter mixture-of-experts model, 46 billion active per token
  • ~1,000,000-token context window
  • Deliberate internal reasoning pass before every answer, thinking time shown
  • Six thinking levels, Light to Hyper Ultra — Hyper Ultra (Plan mode, MAX): draft, self-written rubric, self-grade, revision
  • Dynamic answer length, up to a 9,000-character cap
  • Photo input — up to ten images per message
  • Live web access in Plan mode — up to ten search-and-read rounds per answer, every source cited
  • Image generation — health and anatomy illustrations inline
  • Plan mode progress timeline
  • In NOX WORK: document reading, image background removal and upscaling, connected-app reads

Model details

  • Architecture — Mixture-of-experts — 824 billion parameters total, 46 billion active per token; the 700-billion-parameter 1.3 Astra backbone extended with roughly 100 billion parameters of additional trainable components.
  • Context window — ~1,000,000 tokens
  • Answer length — Dynamic — sized to the question, up to a 9,000-character hard cap
  • Reasoning — Deliberate internal reasoning pass before every answer, with thinking time shown
  • HealthBench Hard — Measured on the 1.4 Orion pipeline: 46.4% (±0.8) at Ultra, September 5, 2026; 53.3% (±1.0) at Hyper Ultra, September 11, 2026 — Nox's own self-runs, official open protocol, all 1,000 examples; not independently verified.
  • Thinking levels — Six — Light, Balanced, Extra, Max, Ultra, Hyper Ultra (Hyper Ultra in Plan mode on MAX, ~5–6 minutes per answer)
  • Raw evidence — Every answer, every grader verdict, aggregate files and SHA-256 checksums are published: Ultra at https://www.getnox.ai/data/healthbench/orion/README.md, Hyper Ultra at https://www.getnox.ai/data/healthbench/orion-hyper-ultra/README.md
  • Photo understanding — Yes — up to ten images per message, read automatically
  • Live web access — Yes — in Plan mode, up to ten search-and-read rounds per answer, every source cited
  • Image generation — Yes — health and anatomy illustrations rendered inline on request
  • Safety layer — Runs behind Leo, Nox's deterministic red-flag detector, on every message
  • Available on — Plan mode on Pro and MAX; selectable in NOX WORK

Availability: Available in Plan mode on Pro and MAX.

Back to the Nox model family