Two of the strongest models released in the last month publish downloadable weights: GLM-5.3 from Z.ai at 753 billion parameters, and DeepSeek V4.1 Flash at 552 billion. Both are mixture-of-experts architectures, both are served through a hosted API as well, and both can be run on your own hardware — if you have enough of it.
What you can download
The licences differ in a way that matters commercially. DeepSeek ships under MIT, which permits modification, self-hosting and redistribution without restriction. Z.ai publishes GLM-5.3 under its own licence; read it before building a product on top.
What it takes to run them
Neither fits on a single consumer card at any quantization. An RTX 5090 holds 32GB; a Q4 quantization of DeepSeek V4.1 Flash needs roughly eleven of them. In practice these are multi-GPU node or Mac Studio Ultra cluster deployments, and the honest framing is that "open weights" here means auditable and self-hostable at datacentre scale, not runnable on a workstation.
If what you want is a model that runs on hardware you own, the constraint to filter on is parameter count, not licence. Our hardware calculator sizes every open-weight model against specific machines and quantizations.
How they score
GLM-5.3 publishes 66.9% on DeepSWE v1.1 on its own launch table, and an independent Terminal-Bench 4.0 run puts it at 41.8% — one of the higher 4.0 scores published, on the harder version of that benchmark. DeepSeek V4.1 Flash publishes 74.2% on DeepSWE and 90.9% on GPQA Diamond but no Terminal-Bench 4.0 result.
Every figure on GLM-5.3's launch table is vendor-reported and has not been re-run independently under a single harness. That is normal for launch-week numbers and worth holding lightly.




