Atria Dawn Preview Review — An Open-Weight 744B Agentic MoE from Shanghai AI Lab
On September 11, 2026, Shanghai AI Lab quietly shipped Atria Dawn Preview: an open-weight, MIT-licensed 744B mixture-of-experts model aimed squarely at long-horizon research and tool-use agents. There was no hosted flagship API and no pricing page — just FP8 and BF16 checkpoints on Hugging Face and ModelScope, a technical report, and a 16-row benchmark table that leads with agent evals rather than the usual reasoning sweep. This review covers what the model actually is, which of its numbers are measured versus estimated on our board, and who it is genuinely for.
TL;DR verdict
| Atria Dawn Preview | |
|---|---|
| Type | Open-weight sparse-MoE agent model (research + tool-use focus) |
| Parameters | 744B mixture-of-experts, post-trained on the GLM-5.2 base |
| Context window | 262,144 tokens (~65K max output) |
| Modalities | Text |
| License | MIT (open weights: FP8 ≈ 756GB, BF16 ≈ 1.5TB) |
| Pricing | No first-party API; open-weight hosted rates (~$0.9 / $2.9 per 1M is a typical third-party estimate) |
| Headline numbers | SWE-bench Pro 59.6 · Terminal-Bench 2.1 78.3 · DeepSearchQA 96.0 · BrowseComp 92.5 · CyberGym 86.5 |
| Availability | Self-host from Hugging Face / ModelScope; third-party hosts route the weights |
| Best for | Long-horizon research and tool-use agents on open weights |
| Caveat | Agentic-eval launch, no standard reasoning suite, and no independent reproduction yet |
If you skip the rest: Atria Dawn Preview is one of the more interesting open-weight agent releases of the month, but it is a preview and its numbers are the lab’s own. Shanghai AI Lab took the same GLM-5.2 base that GLM-5.3 is built on and post-trained it in a different direction — toward research browsing, tool orchestration, and computer-use tasks — then published the benchmarks that direction flatters. The result is a genuinely strong agentic profile at an MIT license, wrapped in the two asterisks that come with every preview: the standard suite is missing, and nobody outside the lab has reproduced the table.
What it is: a second post-training of the GLM-5.2 base
The most useful frame for Atria Dawn Preview is not “a new model” but a second post-training of a base you already know. Shanghai AI Lab built it on the 744B-parameter GLM-5.2 mixture-of-experts foundation — the same base GLM-5.3 is post-trained from — and steered it toward agentic work instead of general chat and coding.
The distinctive piece is the training recipe. Shanghai AI Lab describes a Verifiable Experience Pipeline: every training task runs inside a real execution environment, and only trajectories whose outcomes are confirmed by an external signal — a passing test, a correct retrieval, a completed workflow — are kept. That is the same “verifiable rewards” idea driving most of this year’s agent training, applied at the data-curation layer rather than only in the RL loop. It explains the shape of the results: Atria is strongest exactly where an execution-grounded signal is cheap to compute (research retrieval, tool calls, security tasks) and softer where it is not (broad reasoning, which the launch simply did not measure).
Practically, this is an open-weight model with no vendor endpoint. The MIT license permits commercial use and modification, but the FP8 checkpoint is roughly 756GB and the BF16 is about 1.5TB, so you are either standing up serious local hardware or renting it through a third-party host. There is no first-party per-token price to quote.
What the numbers say
Here is where the case is strongest — and where the asterisks live. Atria’s launch table is agentic and research-heavy, and it is impressive on its own terms:
- Research / browsing: DeepSearchQA 96.0, BrowseComp 92.5
- Engineering: MLE-bench Lite 86.2, SWE-bench Pro 59.6
- Tool use: BFCL v4 77.0, AutomationBench 53.8
- Delivery / computer-use: Workspace-Bench 65.0, JobBench 50.3
- Terminal: Terminal-Bench 2.1 78.3
- Security: CyberGym 86.5
Shanghai AI Lab says Atria posts the highest reported score on five of its sixteen benchmarks and is competitive with frontier agents across the rest. The problem for a cross-model leaderboard is that most of those rows have no column on our board — DeepSearchQA and BrowseComp are not part of the standard reasoning/coding suite. Only two cells map cleanly:
| Benchmark | Atria Dawn Preview | Provenance |
|---|---|---|
| SWE-bench Pro | 59.6 | Vendor-reported (≈ DeepSeek V4 Pro 55.4) |
| Terminal-Bench 2.1 | 78.3 | Vendor-reported (≈ DeepSeek V4 Pro) |
| SWE-bench Verified | ~78 | Estimate — Pro runs ~20 pts below Verified |
| GPQA Diamond | ~90 | Estimate — anchored to the GLM-5.2/5.3 lineage |
| HLE | ~58 | Estimate — anchored below GLM-5.3’s 62.5 |
Read the provenance column carefully. SWE-bench Pro and Terminal-Bench 2.1 are first-party Shanghai AI Lab numbers, recorded on our leaderboard as vendor measurements and scored. But Atria did not publish first-party GPQA Diamond, HLE, or SWE-bench Verified cells — it led with agentic evals instead — so those columns are conservative estimates, greyed out and excluded from every ranking, anchored to the GLM-5.2 base lineage and DeepSeek V4 Pro’s vendor 90.1. They are directionally useful; they are not measurements. With only two mapped vendor cells, Atria sits below our three-cell floor for a full ranking — it shows on the per-benchmark tables it has data for and nowhere else.
The larger caveat is independence. As with most of this month’s launches, none of the Atria numbers has been reproduced by a third party. A model explicitly trained against execution-verified agent trajectories, then evaluated on agent benchmarks, is exactly the case where your own environment is the only score that counts.
What it costs
There is no first-party API and therefore no official token price. As an open-weight model Atria is served at the usual hosted rates by third-party routers and inference providers; our leaderboard shows a typical third-party estimate of roughly $0.9 / $2.9 per million tokens for a 744B open MoE, in line with how GLM-5.3 and other open-weight entries are priced when the lab publishes no first-party rate. The real cost story for most teams is self-hosting: a 756GB FP8 checkpoint is a multi-GPU deployment, so the “free weights” are only free if you already have — or are willing to rent — the hardware to serve them.
Put against the closed agentic frontier — Claude Opus 5 at $5/$25, GPT-6 Astra at $10/$50 — the pitch is control and license, not just price: you get an MIT-licensed agent model you can pin, audit, and run in your own environment. The cost calculator on the leaderboard lets you plug your own token mix and a hosted rate to see where that lands for your workload.
How it compares
Atria’s real peer set is the open-weight agent tier, not the closed reasoning frontier. Against its sibling GLM-5.3 — same GLM-5.2 base, different post-training — it is the research-and-tool-use specialist to GLM-5.3’s broader coding-and-reasoning profile; its published SWE-bench Pro of 59.6 sits just below GLM-5.3, while GLM-5.3 leads on the standard cells vendors reported. Against DeepSeek V4 Pro it is roughly level on the two coding cells that overlap, with a stronger research-agent story and a more permissive license. Against Kimi K3, the other large open-weight release this quarter, the trade is Atria’s agentic-eval breadth versus Kimi’s first-party reasoning numbers and multimodality — Atria is text-only.
Where it clearly gives ground is general software engineering and terminal coding against the closed frontier: the model card itself shows notable gaps versus Claude Opus 5 on those axes. Atria is not trying to win the coding leaderboard; it is trying to win long-horizon research and computer-use agents, and that is where its published numbers concentrate.
Who should care
- Teams building research and browsing agents: This is the headline use case. DeepSearchQA 96.0 and BrowseComp 92.5 are the numbers Shanghai AI Lab wants you to look at, and an MIT license on a 744B agent model is a genuine draw for anyone who needs to self-host a retrieval-and-reasoning loop. See multi-agent pipelines for where a strong research specialist slots into a larger workflow.
- Anyone who needs open weights they can pin and audit: The MIT license and downloadable FP8/BF16 checkpoints make Atria a fit for regulated or air-gapped deployments where a closed API is a non-starter — provided you have the hardware budget.
- Security and computer-use builders: CyberGym 86.5 and the workspace/automation cells point at vulnerability-analysis and desktop-agent workloads. Treat the numbers as a starting hypothesis and benchmark on your own tasks.
- Anyone ranking on broad reasoning or standard coding suites: Wait for third-party GPQA, HLE, and SWE-bench runs before placing Atria against frontier models on those axes. Its published strengths are agentic; its standard-suite cells are estimates. The LLM Benchmark Comparison 2026 covers how to weigh that trade without getting fooled by a vendor sweep.
FAQ
What is Atria Dawn Preview? An open-weight 744B mixture-of-experts agent model released by Shanghai AI Lab on September 11, 2026, post-trained on the GLM-5.2 base. It ships FP8 and BF16 checkpoints under the MIT license, carries a 256K-token context window, and targets long-horizon research and tool-use agents.
Is it open source? The weights are open under MIT, which allows commercial use and modification. There is no first-party hosted API with published pricing; you either self-host the large checkpoints or use a third-party host.
What did it score? An agentic-heavy table — DeepSearchQA 96.0, BrowseComp 92.5, MLE-bench Lite 86.2, SWE-bench Pro 59.6, BFCL v4 77.0, Terminal-Bench 2.1 78.3, CyberGym 86.5, and more. It did not publish first-party GPQA Diamond, HLE, or SWE-bench Verified, so those cells on our leaderboard are conservative estimates.
How does it compare to GLM-5.3? Same GLM-5.2 base, different post-training. Atria is the research-and-tool-use specialist; GLM-5.3 leads on the broad coding and reasoning cells vendors reported. Both are strong open-weight options with different launch emphases.
Have the numbers been verified? Not independently, as of publication. The table is Shanghai AI Lab’s own; benchmark on your own environment before trusting it.
Continue reading
- AI Models Leaderboard — Atria Dawn Preview versus 80+ models on benchmarks, pricing, and context, with a cost calculator and per-cell provenance.
- GLM-5.3 Flash Review — the other post-training of the GLM-5.2 base, with first-party SWE-bench Pro numbers.
- GLM-5.2 Review — the shared base model both Atria and GLM-5.3 are built on.
- DeepSeek V4.1 Flash Review — the budget open-weight agent model in the same value conversation, and how to read a vendor agentic sweep.
- Kimi K3 Review — the other large open-weight release this quarter, with first-party reasoning numbers.
- LLM Benchmark Comparison 2026 — how to weigh agentic benchmarks and unverified vendor tables.
- All Reviews — index of every review on the site.