DE version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
71% Positive
Analyzed from 3884 words in the discussion.
Trending Topics
#model#models#https#glm#alpha#weights#open#com#more#better

Discussion (141 Comments)Read Original on HackerNews
The main issue for me, is that it degraded into a doom loop several times. One of them was running the same bash command about a thousand times. The last model I've used that had this problem was Mimo 2.5, which is quite dated at this point. As a result of this, you cannot leave it unattended / not usable for agents.
See my other comment with examples where 0x Alpha is working non-stop on various projects with zero problems.
The failure mode I run into commonly is agents just stop sometimes. Even sending a "." Or something they start back up, but I haven't worked out exactly how to fix that generally in harness, bit unclear how to tell if they're done or just derped to a stop.
An amusing thought of returning to your workstation to find it as an obsidian block after it gets stuck executing "dd" thousand times.
It sounds like you were using a quant model.
I know that AI is moving fast, but Mimo 2.5 literally came out four months ago. I literally had to double-check after I read this because it felt like just yesterday.
Also, I'm not sure what the industry standard is right now, but my (self-built) harness automatically exits with an error code whenever it detects similar tool calls being sent or when semantic repetition in the reasoning traces reaches a certain threshold. It's pretty easy to set that kind of thing up.
Didnt know they exist - looks very good, maybe even better than Archive.ph
Very impressive model.
Here are some examples, open-source documented and the data available in HF datasets:
https://openzot.github.io/whetstone/ - https://github.com/openzot/whetstone
https://openzot.github.io/arcade/ - https://github.com/openzot/arcade
https://openzot.github.io/machinery/ - https://github.com/openzot/machinery
https://livebench.ai/
while here it outperforms Fable by a significant margin:
https://oxalpha.com/
but if the latter is true, will people still say it was "distilled" from Fable?
It is a capable small model, but it’s not frontier level. The interesting part will be seeing the model size, how it responds to quantization, and how fast it runs on the kind of non-server hardware that we can buy without selling a kidney.
They claim an "independent community benchmark" (pass-fail evaluation on tasks) here: https://oxalpha.com/ox-alpha-vs-fable-5
Source: https://twitterwebviewer.com/?tweet=2091116504787935350
Many people and even software engineers fall for this all the time.
Most of these people are from crypto pivoting to AI doing this.
AI has made this easier and cheaper and it is going to get a LOT worse.
Imagine lots of websites with typosquatting and looking exactly the same as another website, vibe coded and cloned within seconds.
The public have no chance.
The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.)
One thing that I think matters a lot for the non developers, all three of us (sol, oxa, and me) usually agree that oxa's write up is far far better. It explains the situation very well, and has great structure for its write ups. Sol gets the job done, but it's terrible at re-explaining the problem for humans, at laying out information. It also doesn't show it's thinking, so it's imo a terrible peer to work with!
The metric used there is me screaming at my screen per operating hours.
Does it matter? IMO not really. Weights are open after all. (Or.. soon at least for 5.3)
And critically, like contracts in general, Anthropic's terms of service is only binding upon the user/counterparty. So even if a company say specifically sought out 'claude-like' content, and claude code traces available on the internet, if they don't use the Anthropic platform there is no ToS claim.
1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus.
2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe.
I am leaning towards 1.
That's very plausible to have, identify, and fix in a day; especially when you get community feedback in the wild.
Z.AI is the only provider for GLM 5.3 on OpenRouter. I don't see 5.3 on Hugging Face. Not sure if this new model is "full GLM" or something smaller, or if they will like Moonshot AI publish weights but put restrictive license [1], which will again leave Z.AI as single GLM model provider on OpenRouter.
[1] https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE
Related PR: https://github.com/jeffhajewski/latticedb/pull/5
The session used ~100K input tokens, ~60K output tokens, and ~80K thinking tokens.
I reviewed it using gpt-sol-medium, and it seems to be satisfied with it's work.
I gave it an abandoned repo for an Aseprite MCP someone made and told it to iterate with a laundry list of things I wanted from it to include thousands of plugins.
Came back 20 hours later and it shit out a pretty surprising little tool, will post the public repo when I get time.
Seems legit.
It's really hard to know how good it is. So much hype around it.
Chinese labs are not releasing all of their model weights. Qwen is known as an open weight model by most, but their top model is not open weight.
Releasing weights is a marketing strategy for newer labs to get their brand out there.
https://mashable.com/article/chatgpt-gpt-4o-ai-retirement-pr...
> We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing. [1]
People have pointed that this seemingly made Alibaba/Qwen turn around from closing their models (this was rumored after the shakeup early this year [2]) and release the weights for even the Max variant of their new models, which they previously did not.
1: http://english.scio.gov.cn/topnews/2026-07/18/content_118605...
2: https://simonwillison.net/2026/Mar/4/qwen/
https://x.com/davis7/status/2091285712566140986
Wenghi is behind DeepSWE, one of the best benchmarks.
Where? And "Tonight" in which timezone?
Apparently someone working at a 3rd party inference provider also got confused and posted confirmation about it being a glm-flash model, despite having an embargo on that info. Someone jumped in the comments and told them they missed the timezone :)
In any case it should be releasing in a few hours. Timezones are hard.
on toy benches it made quite a few mistakes but was able to fix all of them on its own
(meaning more tokens, more turns, more tool calls — but same outcome as gpt 5.6 sol)
https://x.com/syneryder/status/2091978367579156569/photo/1
Created in a single turn - but technically not a "one-shot", because I gave it a tool to convert SVG to PNG so it could visualize what it had made. I asked it to keep iterating with tools during the same turn until it was happy.
I've also been using Ox Alpha for tasks that better resemble real work, and I'm really enjoying working with it. I've downgraded my Anthropic account so I can put some budget towards Ox Alpha instead, with the rumors that this one is going to be cheap. Opus & Fable are still better at getting large tasks / features done autonomously, but Ox Alpha can work autonomously too, and it's fun. I'm enjoying working with Ox in a way that I'm just not enjoying talking to the 5.0 Anthropic models. (As much as I don't want to say that, as someone with Claude /stickers on their laptop.)
That’s very valid, but right now every other model I use is easier to talk to than Opus 5.0
Opus 5.0 has an impenetrable way of communicating. I can parse it, but it takes so much more work than it should.
As another comparison, I went back to MiniMax M3 for a while last night. It was significantly faster than Ox, but I felt M3's replies were harder to parse, not quite getting to the point. But I guess I could curb that with some prompts.
It depends if the Ox Alpha pricing is as cheap as was being rumored. If it's competitive with DeepSeek Flash and significantly undercutting Luna, that feels like it will be significant.
hard agree. it does not really feel "smart", but the personality is super refreshing
Inference was atrocious in terms of speed and constant timeouts. If it's served fast it will be a delight to use.
Regular pricing is $0.15 input, $0.50 output... but currently 50% off, making it $0.075 input and $0.25 output. That beats most of the V4 Flash providers, but not all, and obviously tokens per task may not be equivalent.
I've also just noticed the blog post reveals the Artificial Analysis score - it's a 57, so it's Opus 4.8 / 5.6 Terra level.
https://z.ai/blog/glm-5.3-flash
It took me a year talking about it until my wife knew that ChatGPT and Gemini are two different things.
PS: some replies, especially if you do a deep dive on comment history, clearly expose the joint effort to drum up support for Chinese models. This has been clear on HN lately as anything even slightly critical of Chinese tech gets downvoted unnaturally quickly. One can just wonder what's behind the effort...
It took me a year talking about it until my wife knew that Kimi K3 and GLM 5.3 are two different things.
Ox is just GLM. And z.ai is the maker of GLM.
The main players in the openweight model market have been known for a while.
And they already have significant user penetration.
But why does that matter? End users (I believe, feel free to correct) do not really contribute all that much revenue-wise. They're certainly not the SOTA target audience.
The professional market doesn't need a household name. They need the most sensible tool for the job, and the CN models right now tick many boxes when it comes to that.
Fast follower persona clusters around emerging zeitgeist across the tellings. At the moment, arguably that's mostly Qwen for everyday hobbyists, and GLM for those that can run 512GB to 1.5TB of memory. This persona is seeking viable applied results: "I have frontier at home".
The early majority pick things up after models are curated into apps like LM Studio or one's platform app of choice, usually at least one major release behind because it takes that long to choose and package into mass distribution.
This is the step where early majority persona "has no idea" what the parade of weird names is about, they care about qualia of the conversations they try to have.
This persona is, at present, very under-served, and likely to remain so until mass devices can perform feeling like 27B at Q4 large quality better, or workplace devices can achieve a pragmatic utility like 135B at Q8 or better.
Harnesses that work where the workplace persona lives bridge this. This persona doesn't care the Chinese model name, they care "does it code?" For that, the applied harness and model take time to be matched, as JetBrains did harnessing a tailored Qwen 3.6 in the IDE. More efforts like https://www.jetbrains.com/junie/ are needed for the majority persona to perceive value from changing their workflow again.
HN's "job" is better outcomes with less friction at each persona.
I would be surprised if the specialist that knows that various Chinese models exist and/or that a user might choose a harness and model separately are a "majority" even of the early variety... in terms of revenue, humans, tokens, or any metric.
(Happy to be proven wrong)
This reminds me a lot of media horse-race reporting, saying that "candidate X has no chance unless they" and "candidate Y has a strong showing in", and it's very thinly cover for the publication liking Y and disliking X, avoiding talking about actual policy, and trying as much as they can to make their predictions self-fulfilling.
Not to go off topic but I am pleased to see open model support from US companies like Poolside.ai, NVIDIA, IBM, Google, etc.
Proper usage in applications will go through similar considerations.
The only "captive" users will be non-tech enterprise users, but they are already in Gemini/Copilot land because they are natural extension to existing Google Cloud/Teams plans and nobody cares about what models are inside those, procurement, contracts and data retention are what matters.
[0]: https://openrating.io/blog/current-state-of-ai-model-fingerp...