Math reasoning researchers
Study a specialized open-weight model and an explicit multi-candidate reasoning pipeline for difficult proof problems.
Independent tool overview
Nomos 1 is a 31-billion-parameter open-weight model specialized from Qwen3-30B-A3B-Thinking for mathematical problem solving and natural-language proof writing.
Visit the official Nomos 1 site ↗
Overview
Nomos 1 is a specialist mathematical reasoning model released by Nous Research in collaboration with Hillclimb AI. It is based on Qwen3-30B-A3B-Thinking-2507 and is intended for difficult problems that require a written proof rather than a short numeric answer.
The model is designed to work with the separate open-source Nomos reasoning harness. That harness generates many candidate solutions in parallel, scores them, consolidates similar conclusions, and uses pairwise comparisons to select a final submission.
Nous Research reports an 87/120 score on the 2025 Putnam problems with Nomos 1 plus the harness, compared with 24/120 for the base Qwen model under the same setup. That result describes a compute-intensive system evaluation, not the quality of a single ordinary chat response.
The weights are available on Hugging Face under Apache 2.0, while the reasoning harness is published on GitHub under MIT. Users can run the model through Transformers, vLLM, or SGLang, but the official examples assume substantial accelerator capacity and recommend eight-way tensor parallelism for the serving frameworks.
Nomos 1 is best treated as a research and experimentation release. Its proofs can still contain hidden gaps, unjustified steps, or confident errors, so important results should be checked by a qualified human or a formal proof system.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Study a specialized open-weight model and an explicit multi-candidate reasoning pipeline for difficult proof problems.
Generate and compare candidate natural-language proofs for competition-style questions with expert evaluation.
Run math inference on controlled infrastructure through Transformers, vLLM, SGLang, or an OpenAI-compatible local endpoint.
Adapt the harness's parallel solving, self-critique, consolidation, and pairwise-selection workflow.
Compare the specialized checkpoint with its Qwen base model while holding the surrounding harness and grading process constant.
Explore generated proof drafts where instructors or experts can carefully verify and annotate every step.
Capabilities
Fine-tunes a Qwen3 thinking model for mathematical problem solving and natural-language proof writing.
Provides downloadable model files on Hugging Face under an Apache 2.0 license.
Pairs the model with an MIT-licensed orchestration system for repeated solving, judging, and final selection.
The harness can launch multiple workers to create several possible solutions for each problem.
Candidate submissions are scored on a seven-point scale until the target number of strong solutions or the time limit is reached.
The harness groups candidate solutions by conclusion and selects a group before final comparison.
Remaining candidates are compared in a single-elimination selection process to produce the final proof.
Official instructions cover Hugging Face Transformers, vLLM, and SGLang.
The harness can call a locally served model through a configurable OpenAI-style API base URL.
The repository includes problems, prompts, runbooks, and generated submissions that help users inspect the evaluation setup.
Process
Step 1
Choose a proof-writing or difficult math task where an open model and substantial inference cost are justified.
Step 2
Check the Apache 2.0 model terms, the MIT harness terms, and any obligations inherited from the base model or deployment stack.
Step 3
Select accelerators and a model format appropriate for a 31B-parameter checkpoint; the official vLLM and SGLang examples use tensor parallelism across eight devices.
Step 4
Load Nomos 1 with Transformers or expose it through vLLM or SGLang at an OpenAI-compatible endpoint.
Step 5
Store each problem as Markdown and preserve all assumptions, definitions, and notation needed for a self-contained proof.
Step 6
Start with one model response and the recommended no-system-prompt setup before adding the more expensive harness.
Step 7
Configure the time limit, concurrency, scoring target, judge model, and prompts, then generate and compare multiple candidates.
Step 8
Check every inference independently, test boundary cases, and send consequential results to a human expert or formal verifier.
Step 9
Save the checkpoint revision, prompts, sampling settings, harness commit, hardware, runtime, judge configuration, and all candidate outputs.
Cost
Nous Research does not sell Nomos 1 as a subscription product. The model weights and reasoning harness are free to download under their respective open licenses, but users pay for the compute, storage, engineering, and any third-party hosting used to run them.
Free to download
Nomos 1 is distributed on Hugging Face under Apache 2.0.
Free and open source
The orchestration code is available on GitHub under the MIT license.
Infrastructure costs vary
Run the model on owned or rented accelerator infrastructure.
Provider-dependent
Use a compatible external host or managed endpoint if one supports the checkpoint.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Science
Consider Gemini Deep Think for a managed frontier system aimed at advanced mathematical and scientific reasoning.
Explore Gemini 3.1 Deep Think →Consumer
Choose Qwen3-Max-Thinking when you want a broader flagship reasoning model rather than a self-hosted math specialist.
Explore Qwen3-Max-Thinking →Business Operations
Consider DeepSeek for a broader open-model ecosystem with accessible reasoning models and hosted chat options.
Explore DeepSeek →Educators
Choose Math Mentor for a simpler conversational tutoring experience instead of operating a research checkpoint and harness.
Explore Math Mentor →Questions
Nomos 1 is a 31B-parameter open-weight model from Nous Research and Hillclimb AI, specialized from Qwen3-30B-A3B-Thinking-2507 for mathematical problem solving and natural-language proof writing.
The model weights are free to download under Apache 2.0, and the reasoning harness is available under MIT. You still pay for hardware, cloud compute, storage, deployment work, and any third-party services.
The downloadable model weights use Apache 2.0 and the harness code uses MIT. Because open-source terminology for AI models can be contested when full training data and recipes are not released, 'open-weight model with an open-source harness' is the most precise description.
Hugging Face lists Nomos 1 at 31 billion parameters. It is based on the Qwen3-30B-A3B-Thinking mixture-of-experts model.
The official model card provides examples for Hugging Face Transformers, vLLM, and SGLang. The reasoning harness expects an OpenAI-compatible API endpoint and can point to a locally served instance.
It is a separate orchestration system that creates candidate proofs in parallel, scores them, consolidates candidates by conclusion, and uses pairwise comparisons to select a final submission.
Nous Research reports 87/120 with the Nomos reasoning harness and human expert grading. The base Qwen checkpoint scored 24/120 under the same conditions.
No. The reported system uses a multi-candidate harness with concurrency, repeated scoring, consolidation, pairwise selection, and a time budget, so it should not be interpreted as ordinary single-pass accuracy.
No. A fluent proof can contain a subtle invalid step. Verify every conclusion independently and use qualified human review or formal proof tools for important work.
Bottom line
Nomos 1 is a useful open research package for teams studying difficult natural-language math reasoning, especially because both the checkpoint and its multi-candidate harness are inspectable. Its headline Putnam result is interesting but comes from an expensive system-level workflow, not a single response. Use it when openness and experimentation outweigh ease of use, and budget for substantial inference plus rigorous proof verification.
Visit Nomos 1 website ↗
SciSpace BioMed Agent - A specialized biomedical research co-scientist for multi-omics, clinical phenotypes, and lab protocol reasoning

TRIBE v2 - Meta's predictive foundation model that simulates human brain responses to sights, sounds, and language

Edison Analysis - Next-gen scientific analysis agent

Claude Science - Anthropic’s app for AI-powered scientific research

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.