FR version is available. Content is displayed in original English for accuracy.
Advertisement
Advertisement
⚡ Community Insights
Discussion Sentiment
70% Positive
Analyzed from 9248 words in the discussion.
Trending Topics
#models#model#flash#glm#don#more#chinese#same#https#opus

Discussion (360 Comments)Read Original on HackerNews
I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.
I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.
I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.
https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...
Amazing results for open weight and that size, but a really long way off, and I'm extremely skeptical of benchmarks that show these smaller models as being anywhere close to Opus 4.8 (or even earlier Opus's).
DS4 Flash 0731, on the other hand, wildly opposite experience. Would recommend.
GLM 5.2 - even quanted down to a hybrid 4/3 bit setup is amazing for everything but the hardest/most complex stuff in the same projects/realm.
The tradeoff is time (especially on RDMA4 hardware) - it does take a long time and spend a lot of tokens to get to the result, but I've found I can trust the results enough that I can queue a lot of work, essentially have it running all the time and achieve a decent velocity.
It's the first small local model I've felt like I can do real work with.
Wow, if you don't mind me asking. How and where?
They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.
$4,000 isn't priced insanely? ye gads
Reality is on a single spark I'm constantly running out of room and it being an odd size M.2 slot it's a pain to upgrade. I'm setting up a NAS over RDMA via ConnectX though, that's fun.
Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly.
There are a lot of social media posts about people cancelling their Anthropic or ChatGPT subscriptions after installing a local LLM. I’ve used local LLMs a lot and I spend a lot of time with frontier models and the difference is still huge. As far as I can tell, the social media posts about local LLMs replacing frontier models are either wishful thinking, engagement bait, or people who must be working on much simpler projects with a much higher tolerance for slop than I have.
Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it
The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled to how differentiated their marginal contribution is to peers. They’re much more incentivized to spend their personal/work time optimizing for being more skilled or acquiring some kind of competitive advantage relative to baseline.
Most people don’t consciously run the numbers of “I get paid $X/hr to add $Y of value” or model pay at work as something with variable inputs (eg something that can be increased with high performance), so it makes sense to them to spend 20 hours of time to save $100 or to make themselves 5% less efficient to take home 0.5% more or avoid doing something they don’t want to start doing.
NOT saying this always happens or that they’re stupid for doing so. I didn’t even realize how much I had been doing it myself until I started recognizing it, and shifted to having my own comp/performance fully aligned with the company’s P/L.
It actually makes a lot of sense IF you can accurately estimate incremental upside (which is much harder and more diffuse than modeling downside if you’re salaried a employee) or if the upfront skill/knowledge investment that looks like bikeshedding pays off in the long run.
Don't use DS4 Flash in max effort mode. It's just spinning its wheels, in my experience (I have a harness for testing models with 25 real bugs/features/etc from my real projects that I measure outcomes against) DS4 flash does _worse_ with max effort. It will literally have the right approach and reason itself away from it.
The more reasonable comparison is against rented GPU's, while looking at tradeoffs in latency and upload/download/storage/instance management overhead.
Buying hardware for local models is meeting a wholly different need than buying tokens through OpenRouter or whatever.
For example, I have a small posix-shell-based LLM harness that can SSH into my NAS and run organization tasks using the local DS4Flash that I have right now. It's already been a massive help for me to keep me organized, and that's just 2x DGX Spark's worth of compute.
But things change real fast when you're no longer bound by costs/apis/rate limits. All of a sudden it's not about "how can I do this right and efficiently" and more about "I can poke at and test _all the things_ that might make this better".
I think most people who can't see this value in the local inference approach are likely still copy/pasting from their web LLM ui's or don't even come close to subscription quotas. Meanwhile, 1b tokens a day is a light day for me with 3 $200/m subscriptions + some level of sub at basically every frontier level provider. Had I been less frugal and ponied up for the hardware before things got crazy I wouldn't need 80% of that - just the frontier models for the most complex tasks, the open weight models would handle the rest easily _and_ I'd get to do a lot more exploratory work without concern about quotas.
https://deepswe.datacurve.ai/
That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost. Roughly equivalent to sol medium, at a fraction the cost.
They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts.
Congrats to them!
Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium tends to be rated pretty poor on agentic tasks compared to the Chinese lab open models, but I find the latter much more likely to lose track of their own behaviour during a long-horizon task or get stuck in a doom loop.
(This isn't a comment on GLM-5.3 Flash as I've not used it!)
https://kommodo.ai/i/IFSUUQYT522uZXWvePZ1
it still fells stupid sometimes and it is benchmaxxed for sure. but its good enough that im building all the hobby projects with it.
I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a single line of code.
https://www.reddit.com/r/codex/comments/1vj3hhn/wait_so_sol_...
https://www.reddit.com/r/codex/comments/1vp0rig/sol_can_fina...
I’m pretty sure that plain Sol, serially, could have finished the task faster, cheaper, and far more accurately. I’m also pretty sure that any competent subagent orchestration could have gotten it done with even very simple subagents quickly and cheaply.
(Is it really that hard to set up a handful of subagents that all use the same initial context and to load that context with what actually matters? The APIs certainly support it.)
The flash one?
My prompt was something like: "here's data I have, here's what matters to me, create HTML mockup".
All GPT 5.6 models were laughably bad. And I don't want to downplay it - they were just absolutely, objectively horrible. Every single attempt was what I could probably call "if json was ui".
Claude models produced... "claude look".
GLM 5.3 - somewhere between GPT and Claude.
Kimi k3 - each attempt produced beautiful UIs. It used components that I didn't even know existed and wouldn't even know to ask for. But expensive, very expensive.
ox-alpha (GLM 5.3 flash) was very close to K3. And at this price point, it's already configured as "designer" model in my oh-my-pi.
It's what people know. Opus is just the common target.
> Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash
The problem with this and DeepSWE is it goes for a very specific profile. I'm not convinced DeepSWE is any accurate in actual work. It's surely a different signal (compared to some that allow cheating) but it has its own issues, e.g. weak harness.
Luna is great at following instructions but bad instructions or anything not covered = death.
Deepseek is more analytical. Good for bug tracking.
GLM is a better all rounder in some ways. Better at creativity.
July 16th: The "Kimi K3 moment" - China has caught up to Opus!
4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third!
12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!
I'm sure they have nothing to rival this on a price/performance basis and have already given up on that
Broad and perpetual license over inputs and outputs, and even your name and profile picture.
Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country.
Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.
Vague prohibitions on discussing Z.ai, even my posting this comment violates it.
Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan kiss your money goodbye.
Nearly every TOS I've ever read has a "We can ban you for any reason, or no reason, are under no obligation to disclose any reason." line somewhere in it.
HN's for example
> We reserve the right, at our sole discretion, to change or modify portions of these Terms of Use at any time.
> You acknowledge that Y Combinator may establish general practices and limits concerning use of the Site,
> You further acknowledge that Y Combinator reserves the right to change these general practices and limits at any time, in its sole discretion, with or without notice.
> Y Combinator reserves the right to investigate and take appropriate legal action against anyone who, in Y Combinator’s sole discretion, violates this provision, including without limitation, removing the offending content from the Site, suspending or terminating the account of such violators and reporting you to the law enforcement authorities.
Not even close. Even OpenAI and Anthropic aren't bad enough that they claim literal ownership of your inputs and outputs.
> HN's for example
You're not paying to use HN. Getting banned here has essentially zero consequences.
If Z.ai uses its absolute powers to ban you because you wrote a review about them or something, then you lose actual money. This is especially relevant if you're looking to take advantage of their discounted yearly payment option.
At most I suspect the A.I. providers will just come up with yellow banners like Anthropic did where naughty smut writers get put in the time out corner.
OpenAI revoked my Cyber verification, along with many others, asked to reverify (i.e. give my biometric information to Persona), had me do it 8 times, just to find out several days later that they silently implemented a nationality whitelist, and my nationality didn't make it (and no, it's not a sanctioned country).
Their support says they can't look into anything or do anything, and their public spokespersons on X deny everything.
I get tons of cyber refusals now (lots of reverse engineering), so it's only matter of time when my account is going to get banned.
At least Z.AI is being honest here. And no provider other than OAI/ANT had me submit my biometric information just to use Ghidra.
How did you discover this?
I opened the Persona tab once, closed it and the tab never opened ever again. "Precheck failed".
What countries are banned? I'm from Brazil.
I went as far as initiating an LGPD (brazilian GDPR) process against them due to this. At some point I got it in writing that I'm allowed to make a new account and try again. Until now I was assuming it was just some weird account state. If I'm banned from TAC due to my nationality that's seriously disgusting...
A week later, I tried testing the TAC flow on my SO's account, which had never had a TAC attempt before. Selecting Georgia in the Persona iframe now says, "We are unable to verify identities in this country."
But I know this isn't a Persona limitation, as I verified with Anthropic using Persona the same day.
So what I think happened was this: OpenAI silently implemented a country whitelist on their end and revoked TAC for affected individuals who already had it, calling it a "technical issue". They forgot to disable those countries in Persona, so everyone just got a cryptic error. Then they disabled them in Persona too.
Interaction with support was AI with human names, which essentially just repeats what you said. And it ended with:
> I’m unable to provide additional details about verification outcomes, and Support cannot manually override the result. At this time, Trusted Access for Cyber verification does not support retries or appeals.
I even provided them my credentials, and support AI was basically: lol wat we're here to check technical errors, your credentials are of no relevance.
Then alternatives are:
- Grok - where I absolutely have 0 trust in X.ai's interst in "pushing humanity forward".
- OpenAI and Anthropic - which seem to try to be building the biggest moat they can by pushing to ban open models. And at the same time want to be an Arbiter of what level of intelligence I can use.
- Google and Meta - I don't need to talk about the practices of these companies.
Yes, the terms of service aren't great. But the alternatives aren't great either. I don't believe that a future which OpenAI and Anthropic are pushing for has my best interest in mind.
I don't believe in that either, but these totalitarian terms are absolutely unacceptable.
Chinese companies do not follow American laws and there are absolutely no consequences for violating it.
Moreover, the average American is not even aware of exactly what the legal/judicial environment is like in China. If your code and data is stolen, you can't fly to China and demand justice in the courts.
... lmk when anthropic/openai/spacex/xai are held accountable for anything. Anything at all. Hard to be when you're _writing_ the rules.
No, they don’t. This is an absurd statement to make in 2026.
In any case what matters is what is enforced in practice. It will be a mild inconvenience to switch providers on Openrouter.
If Anthropic or OpenAI decide to apply those same arbitrary terms, you are SOL.
Yeah, I've compared both. The US companies generally aren't as vague, and they don't claim ownership over inputs and outputs.
> They also don't require persona id verification, witch is wat turned me away from openai.
Could be worse. I was dumb enough to verify, only to get rejected for unknown reasons with no retries and no appeals. Had my privacy violated and have nothing to show for it.
Yeah, running frontier open weight models on my own hardware has essentially become my dream at this point. I hope the hardware manufacturers step up production to meet consumer demand.
>Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country.
[...] may cause harm to Anthropic, our users, or third parties, we reserve the right to remove or take down some or all of such Third-Party Content using, where appropriate, algorithmic and human review.
You may not export or provide access to the Services into any U.S. embargoed countries or to anyone on (i) the U.S. Treasury Department’s list of Specially Designated Nationals, (ii) any other restricted party lists identified by the Office of Foreign Asset Control, (iii) the U.S. Department of Commerce Denied Persons List or Entity List, or (iv) any other restricted party lists
>Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.
we will use Materials for model training when [...] your Materials are flagged for safety review to improve our ability to detect harmful content, enforce our policies, or advance our safety research.
>Vague prohibitions on discussing Z.ai, even my posting this comment violates it.
>Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan kiss your money goodbye.
To engage in any other conduct that restricts or inhibits any person from using or enjoying our Services, or that we reasonably consider exposes us—or any of our users, affiliates, or any other third party—to any liability, damages, or detriment of any type, including reputational harms.
Mind you, that's Anthropic's Terms of Use in Europe. I have zero doubts the TOS applied to the US is even worse and that merely mentioning your first born in a chat entitles them to a part of its soul.
I have prompted out a lot of disturbing and inappropriate content with GLM-5.2, that would have left other American models blanched in the face or clutch their pearls. I think this is mostly a reference to Anti-CCP stuff.
In fact, I don't think I've ever even had a prompt refused.
Just like that we are witnessing an open burial. It's now in everyone's interest to keep the valuations in the 'A.I' economy as they're though it's apparent they're not justified.
whether it's the cost to develop models, cost of hardware, cost of serving ie inference.
Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them.
I highly suspect that the Gemini Google uses internally is very different from what they offer in Antigravity.
A fun test would be to compare the logits for "gemini", "claude", etc for a continuation of "I am " on all these models. I'm sure that e.g. GLM, Qwen, DS are dominant, but I'd be curious to see the next highest contenders, and how they compare to each other.
Gemini 3.7 Flash is a pretty great model IMO. You shouldn't compare it to Opus, Sol, K3, etc since it's a much smaller model but it's a little better compared to Sonnet, Luna or Terra, etc.
RIP Nivida shareholders
Further quote:
"Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."
https://z.ai/blog/glm-5.3-flash
https://en.wikipedia.org/wiki/HiSilicon#Ascend_910
https://medium.com/@huaweiclouddevelper/a-brief-introduction...
And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet.
So that leaves local hosting/leasing, but one of those has totally non-practical economics and the other doesn't have enough compute to meet any kind of real demand.
I also have yet to meet a single person who isn't neck-deep in the tech space mention a Chinese LLM. It's 100% the big American three.
If anything it's custom chips from the labs that threatens Nvidia.
https://www.ft.com/content/32a70a3c-7d28-40b4-808e-36edb58c7...
Because I genuinely can't tell if you mean Google or SpaceX/X.ai lol.
xAI is already selling spare compute, and basically exists just to gas spacex's perceived valuation.
While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.
That doesn't mean by any stretch of the imagination Nvidia will disappear. But the entire stock market valuation, not just tech, has had me scratching my head for a while.
Curiously, there is not a single real CUDA competitor anywhere in the world. We almost had one with OpenCL, but all of the American stakeholders abandoned it right before the crypto/AI takeoff. All of which means that Nvidia sets their own margins, exploiting American investors and taxpayers while letting China avoid their dominance. So the American economy subsumes the bulk of Nvidia's arbitrarily-priced debt, and the Chinese economy can direct SOEs to pour billions in liquid cash into real GPGPU research.
I'm an American and I'm pretty fond of Nvidia, but Jensen was right about this policy; it gives China everything they need to actually replace CUDA. It's reminiscent of America's attempts to deprive China of ARM and Texas Instruments IP, only to end up swimming in unlicensed clones after refusing to sign an IP deal.
This isn't the kind of thing you can hash out in public and go back and forth on. Once you put it out there, the other party will take steps to make sure they don't have to rely on us in the long run.
Zai is on another "export control" list outside the broader 1. Doesn't help.
RAM was probably the bottleneck for the amount of context they were offering.
I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful
... I'm at a loss for words here. It was being served for free. To the entire world.
> ... I'm at a loss for words here
No need to be so dramatic. I think it's great that they're developing chips, but the whole "RIP nVidia" claim was overly dramatic.
Edit: Ah:
> This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.
> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.
Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)
Compared to the rest of the world?
I'm sure the chips are fine, but they clearly didn't have enough capacity for the demand they had (that 100T/day claim was asbolute bs)
Boy, do you have a rude awakening in store.
Get a 210 strike put contract and if your thesis is that nvidias current 10 day slide continues you could make some money.
>> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."
It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.
[0] https://news.ycombinator.com/item?id=49397204
[1] https://news.ycombinator.com/item?id=49431231
This announcement doesn't really mean anything at all. It means the very few people who are already using Z.ai's API will continue to do so, but the vast majority of money going to Nvidia is through the massive amount of business going to Anthropic, OpenAI, and other western cloud providers and inference providers, who are mostly using NVidia chips for inference.
Also, NVidia chips are still sold out and supply constrained.
Now it's similar cost to DeepSeek v4 flash, but smarter.
My tests: https://aibenchy.com/compare/z-ai-glm-5-3-flash-max/deepseek...
This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home.
@edit: so many releases that I forgot to math. This fits just fine in q4, realistically the minimal hardware would be 192gb - so blazing fast on double rtx 6000 pro and usable on 256gb unified memory. You could even go with 5bit quant on 256gb.
… you’ll still need to splurge, though.
https://github.com/JustVugg/colibri
Generally, for local consumer use, these large MOE models are best for unified RAM systems like DGX Spark or Mac Studio.
Artificial Analysis ranks GPT 5.6 Luna similar to GPT 5.4, but that never matches my real world experience. AA seems to do a good job making a single number as representative as possible but there is still so much benchmarks don't communicate.
For implementation tasks, where I have the problem already defined and researched, or just simple task, I'd definitely use something like Luna xhigh or max. If the task is vague, or involves planning, I'd rather use Sol medium, even though it's theoretically worse on benchmarks.
From a biased source, but would be big if true. I've had great results with GLM 5.2.
From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.
It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.
What irks me about this is that the harnesses seem to be just an afterthought here.
Don't get me wrong, I love messing around with installing Pi, getting it hooked up with OpenRouter, and just trying all kinds of different stuff, local models, etc... but when it comes to literally just setting up a productivity environment and trusting my entire machine with it, I just run Codex.
I have heard from anecdotes where people have indeed replaced their main drivers with DeepSek V4 Flash or GLM and state that "it's almost as good as... [claude/gpt]" but I never hear anyone say "yeah, this is the model/harness that I now run on my machine and don't mess with it"
Have been using it as my primary harness for personal work for I'd say 6 months. I recommend everyone create their own harness at least to learn. There are a lot of practical benefits.
* me raises hand.-
Think z code gives a token bonus though
I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.
FWIW, people were saying "ASICs will kill CUDA demand!" since the crypto mining boom. Then a few months later, CUDA found another niche application in LLM applications.
With the mounting demand for robotics, surveillance and autonomous weapons, I don't see how Nvidia couldn't keep their spot. They have their pick of the litter with hundreds of market segments, and unlike the rest of FAANG they're not afraid to branch out.
│ https://openrouter.ai/api/v1/chat/completions model: stealth/ox-alpha auth: OPENROUTER_API_KEY status: 404 Not Found response: {"error":{"message":"Thank you for participating in the Stealth Ox Alpha testing period. This model was ZAI's GLM-5.3 Flash.
│ Use it now: https://openrouter.ai/z-ai/glm-5.3-flash","code":404},"user_...":"}
With GLM 5.1 and 5.2, the big problem was tool calling and long-horizon coherency. 5.3 was more trustworthy at the cost of longer thinking traces, and now Flash seems to improve on it once again with a more concise, smaller model. As long as there aren't any noticeable regressions, I could see myself defaulting to this for >90% of my day-to-day coding work.
How is the business model of Anthropic/OpenAI will sustain?
e.g. "Agent Coding Performance by Effort Level" cuts Y-axis from 0~20.
- This makes it as if GLM-5.3-Flash made a bigger jump than it claimed as the Y-axis does not increase much (stupid trick used in biz reports)
I did mention that ox was working ok for me, and having an open-weight comparable to close to SOTA makes it very compelling for me to try it out locally (well, only if I got more VRAM)
edit: nevermind. it is there in the artifical analysis scatter plot, but is greyed-out.
MUCH more interesting is that in that chart, their cost is WAY off. The actual chart shows GLM 5.3 Flash at $0.09, but their chart shows $0.045...
Qwen 3.8 Next Flash: 125B + 51B = 176B parameters with 6B activated
DeepSeek V4 Flash: 284B with 13B activated
The new Qwen model is the most promising for one or two Strix Halo 128GB with the low number of active parameters. On paper it's much stronger than Qwen 3.8 27B.
- Input: $0.15 - Output: $0.50 - Cached input: $0.03
https://openrouter.ai/compare/deepseek/deepseek-v4-flash-073...
EDIT: Looks like they are swizzling around the pricing dynamically on that page, on both the GLM and the DS sides, so who knows.
What I'm saying is, if you're expecting a model that can be run on a 16GB or 32GB machine with the intelligence/knowledge of Mythos or Sol, it will never happen. It cannot happen, just like you cannot watch the Odyssey saved as a 16MB file.
Smaller models can get faster and smarter, but by definition they can never compress all of the knowledge of a frontier model and they will approach a limit by which they cannot get better.
https://www.deviantart.com/sssfjknfvdknj/art/the-ENTIRE-shre...
The ‘frontier’ models rely on scale to achieve their results but that’s not the only approach. Eventually we will hit up against the fundamental limits but we are not close with Sol and Mythos.
Like were executive at Google and AWS and Microsoft expecting this kind of performance from models smaller than what openai/anthropic have been doing? Are we really in a "compute desert"?
Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?
It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.
That's good. Keep going.
Now the US is behind in EVs can you guess what they're doing? [1]
[1] https://evwire.com/p/video-ford-ceo-jim-farley-says-they-fly...
If the USA wanted a copyright treaty with China bad enough, we would negotiate one. China is not breaking any laws here, international or otherwise.
Intellectual property is part of WTO agreements but enforcement is domestic.
US companies do it too, regularly, they simply hire and poach staff from competitors.
Proving it to be IP theft is difficult unless you can prove documents being passed. But often all you need is the know-how of the hired talent.
"problem" indeed.
(281 points, 118 comments)
> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.
It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.
https://x.com/Zai_org/status/2092616204787626030/photo/1
Give it a day or two. More will pop up.