Hacker Newsnew | past | comments | ask | show | jobs | submit | batperson's commentslogin

The future of inference is likely in ASICs, so we'll get the inverse, a bit less capable than frontier but super fast models. Like this 14k tok/s beast https://chatjimmy.ai/ from Taalas (who got acquired by AMD recently).

GPT-6-astra runs at like ~40 tok/s, I have a hard time imagining what could be accomplished with that type of model at 10k+ tok/s when in the hands of the public. Will certainly make cybersecurity a challenge for older systems.


This is incredible! Are there other big players in this space (freezing models to silicon)?

take a look at Cerebras, who are doing wafer-scale compute

I imagine Astra is/will soon will be on Cerebras?

Nope, OpenAI partnered with Broadcom to produce their own chips and the performance/watt looks good

Is it definitely out of the question?

https://openai.com/index/cerebras-partnership/


Like how crypto used ASICS but then didn't because the scaling of consumer hardware made it obsolete?

To the extent that cryptocurrency moved off ASICs, it was because of interest shifting to different cryptocurrencies that were specifically designed to be harder to mine on an ASIC than Bitcoin's compute-heavy, memory-light hashing.

I'm not sure there's any reason to expect a similar shift from LLMs. The hardware used for training doesn't dictate what hardware needs to be used for inference, and nobody's going to design an LLM architecture with an overt intention to make it better suited to GPUs and hard to target with ASICs.


Yet it doesn't seem that ASICs will have any particular advantage over consumer hardware since AI is very memory heavy, which is (right now) expensive no matter how you package it. And the compute is just simple matrix multiplication, which is almost entirely what GPUs were meant to do anyway.

ASIC vs GPU doesn't make a ton of difference when both are relying on commodity DRAM; in that sense, LLMs are more like the anti-ASIC cryptocurrencies. But the actually interesting ASICs are the ones that ditch the commodity discrete DRAM chips. They lose out on the memory density and thus struggle to scale up to the largest models, but for what does fit onto a Cerebras wafer or a Taalas chip, the speed is phenomenal. They have a real shot at securing the "smart enough, and really fast" segment of the market.

And it seems more plausible to me that an ASIC architecture rather than GPUs would be able to best make use of something like wafer-bonded custom memory to approach the density of discrete DRAM while retaining the extremely high bandwidth that comes with arbitrarily wide interfaces and minimal PHYs.


Go back and correct your idea that consumer hardware made asics obsolete. Then we can figure out if asic or asic like devices for inference will have no advantage.

Except Taalas is much faster than GPUs, orders of magnitude so. They aren’t going to get 100x faster at inference any time soon!

There's a new SOTA model every few months, are you supposed to buy a new chip every new release?

Yeah! Nobody needs chatjimmy.ai. Nobody needs their results to come back instantly instead of at 10 tokens per second. Nobody needs a CPU faster than a megahertz.

This is factually wrong no? Bitcoin is asic only. The others all changed for other reasons unrelated to your thought.

My thought was that ASICs turned out not to be worth it for crypto mining because consumer hardware evolved fast enough to do it, while also being cheaper and having some resale value, while ASICs are useless besides mining and have no resale value.

So I'm extrapolating this same idea to LLM inference.


For the problems ASICs exist they vastly outperform general hardware. Typically both in absolute speed and efficiency.

But it's only possible to make custom ASICs when you have a specific problem to solve. For newer crypto systems they can vary enough parameters that building a flexible enough ASIC to recoup the investment before the algorithm changes and makes your hardware useless.

For problems where the problem to solve remain in the problem space the ASIC can solve there is no point to use a thing else.


You’re extrapolating on something that is false. Consumer hardware never caught up to asic.

x86 has a built in instruction for doing AES. That's just moving the ASIC into the CPU core, not eliminating it.

Is there any cryptocurrency that uses AES?

I hate the word crypto, very ambiguous. In my professional life it almost always refers to cryptography.

Am I missing some joke here?

Do you know what reasoning level this was generated at?

The video compression created the illusion, it removed enough frames to where you can't tell the motion of the spokes since every frame refresh they moved enough distance to be "valid" for both forward and backward motion.

I had 3x32" 4k at 100% scaling for the longest time (32" 4k is still the best bang for buck right now imo). Main horizontal center display and a vertical on each side, angled aggressively towards my head.

But 4k is quite weird for productivity, if you do two columns you get 1920px on each side with a split right in the middle, which is far from ideal. With 3 column layout you end up with 1280px a piece, which ends up being too cramped with most apps.

Recently I got a 40" 5120x2160 LG display which ends up being about the same PPI as 32" 4k, but it lets you have three 1706px columns, which is much more comfortable. It's quite nice, but unreasonably expensive. I could've bought three 32" 4k for the same price.


Buy what? Hardware for inference? Pretty sure you'd still be giving money to "AI assholes", just somewhat different flavor. And you'd likely have to pay so much that the $60/mo would seem like peanuts, and in the end you'd still have a subpar experience/performance compared to SOTA.


I own an elgato Stream Deck (somewhere in a drawer), I love the concept of keys being a display but the keys are VERY mushy. Still a better deal and a way more versatile device than that Codex Micro pad.

Now that I think about it, I think I'd enjoy using streamdeck more if it was just a USB touchscreen thing maybe with some vibration for tactile feel with the same UI.


Social people will be fine, I think this tech is far more important for lonely people who for any reason don't get to socialize much (if at all), this is especially common in older people. These people might not have any other alternatives.


I think a fundamental and flawed assumption in your general line of argument boils down to that line:

> Social people will be fine

Thing is, "social" is not only a scale, it's not even a line for scale. It's this multidimensional space of different social preferences and different preferences as to _how_ to socialise, not to mention with whom and for what reason. In short, it's much more complicated than as to permit a "binary" division into the "social" and "other" people, but even if you did, I would wager the division puts it at 50% for either group, meaning you're basically describing a cataclysmic event where 50% of the population suddenly has no access to a fundamental trait of humanity that in known and unknown ways has assured our feeling of happiness and warded off a plethora of conditions of misery and worse that psychology has no names for (and won't ever have names for because that's like classifying different gradations of drinkable water, I imagine).

But yeah, humanity has always had to evolve, it's just that these changes come too fast for our adaptability to adapt to, I believe.


> I think this tech is far more important for lonely people who for any reason don't get to socialize much (if at all), this is especially common in older people.

Uhm, those lonely people need to get out and start talking. How is this going to help society? This is going to make it worse.

Oh that kid is kinda quite and sad? Throw him an iPad. Oh that adult is kinda bored and wandering aimlessly? Throw him into a casino. Oh that adult is kinda lonely and feels like they don't have anyone they can talk to about their life? Give them LLM companionship.

Yep, it is over for humanity. People simply don't understand externalities.


What about people with disabilities, physical or mental, who can't get out, old people with no family or family who doesn't care? In theory it would be great if everyone got some attention and socialized, in reality a lot of people in society are forgotten, no one wants to talk to them.

For some of these people even talking to a robot would likely be a huge improvement in quality of life, and that's just talking. If said robots could also help them out in real life, sort of like a personal assistant, that would be even better.


These people are an opportunity for the rest of us to be better people, to show compassion and mercy. They are not worthless to the rest of society. Giving them LLM companions says that they are — that their only self-worth is their feelings.


Who’s paying for this? Nice sentiment, but last I checked prices weren’t coming down, and we’re talking about a cohort that doesn’t generally have a ton of extra cash.


> What about people with disabilities, physical or mental, who can't get out, old people with no family or family who doesn't care? In theory it would be great if everyone got some attention and socialized, in reality a lot of people in society are forgotten, no one wants to talk to them.

I mentioned in another comment that I totally get it for these people. You are using the extreme cases to justify opioid for the masses. The moment you "solved" human loneliness and interacting by giving them drugged up with chatbots, its over. I am telling you, it is over.

AI psychosis is bad enough as it is with really poor imitation. The moment it can imitate and fool the average person, I just don't see why most people need real friends anymore. We will NEVER be able to get it out of kids hand, the cats out of the bag. You think kids growing up with social media and ipad is bad? Lets see what happens if they live in actual fantasy land with their robotic friends.

I don't even know if I want to live in that world anymore.

Also, your solution to lonely people with disabilities or old age is to create robotic "Friends" for them? This is just sickening. You might as well just drug them. The _actual_ solution is to create a compassionate society where this doesn't happen. The solution to overworked society isn't more opium to relieve the pain.


>Uhm, those lonely people need to get out and start talking.

Great, what is being done to help that happen?


Unironically just "man up". I get that there are some people that have actual sickness that prevents them from socializing but your little anxiety does not count. Believe me, I know I have the same problem. I quite literally had no friends for almost 7 years or after highschool. Society can't afford to babysit a 30 years old man that have anxiety and doesn't want to put any effort. My parent had to call me to check if I am alive, and even then, I don't feel like talking to them after months of no talking. I had zero interest to form any kind of relationship. That feeling that you have when you are feeling like "I would rather just order Uber", yea, you need to OVERCOME that. That is the effort. Ruminating, thinking and fantasizing about what "could have been" does not count.

Even if all I am saying is bullshit, "what is being done to help"? What about what is being done to make it way, way, way, way, way worse. If I had LLM companionship when I was alone, yea, I would have never gotten out of my shell. I would be stuck in there forever. Hell, why should I even talk to you? I should just argue with a robot instead.


AI girlfriends, apparently.


Edge user here. For one, chromium is faster than firefox, any given page will load about 20% faster, another reason is edge workspaces feature, I've grown to like it, which seems to be some sort of chromium feature that everyone bakes in weird ways if at all, and I'm still running ublock origin on edge without any funky bypasses.

Then there's a fact that a bunch of sites/webapps straight up refuse to work on firefox and they ask you to install chrome or something. And lastly chromium the most popular browser flavor and as a web dev it helps to see pages through "the same eyes" as my users/customers.

That's about it, the only reason I use firefox every day is their superior picture-in-picture player, chromium one is waaay inferior.


> Edge user here. For one, chromium is faster than firefox, any given page will load about 20% faster,

I'm skeptical; You're probably measuring Chromium + ads against FF + ads.

The only fair test is testing agains FF + uBlockOrigin. And there, FF wins hands down.


I'm hardcore FF, and it used to be a bit slower than Chrome, but nowadays the difference is barely noticeable. And on very large pages (e.g. big tables), Chrome is a lot slower than FF.


To access Edge Workspaces, you’ll need a desktop running Windows 10, Windows 11, or Mac OS, Microsoft Edge version 144 or later, and to be signed into Microsoft Edge with a Microsoft (MSA) account or Microsoft Entra ID / Azure Active Directory (AAD) account.[1]

> Then there's a fact that a bunch of sites/webapps straight up refuse to work on firefox and they ask you to install chrome or something.

This is rare in my experience. And most were fixed with an extension to change the user agent string. Or were for amusement and used a new Chrome feature. Or used a feature Mozilla rejected for security and there were alternatives.

[1] https://www.microsoft.com/en-us/edge/features/workspaces?for...


Give Vivaldi a try. I used edge on windows and android ever since it started used chromium, and switched to Vivaldi on Linux and android 8 months ago. Generally quite happy with it - not really missing any features from edge.


Yeah, I am still salty that Firefox removed the old Tab Groups feature. The new one they reintroduced in 2025 is pale in comparison.


If you check openrouter there are a tons of providers selling API access to open source LLMs at a fraction of the cost compared to SOTA models (codex/claude). What model you're serving and what kind of platform you serve is a big factor.

I'm no expert but I think eventually we'll have even more specialized ASIC like machines with models burned into them and a that will absorb a chunk of the market, similar to what happened to crypto mining but to a lesser degree since the work isn't as static.


NN-specific ASICs won't buy you much more FLOPs per watt than GPUs/TPUs will. These chips are already extremely good at NN computation. Sure, you could remove GP shader support and free up 5% of your die for a few more cores (which btw is what TPUs pretty much are), but that's about it.

Either way, you'll still be starving for data.

The best work in this area is memory-integrated Big-Ass-Die or Big-Ass-Chiplet solutions like Cerebras which park SRAM right next to your cores, not ASICs.


>but I think eventually we'll have even more specialized ASIC like machines with models burned into them

This has already happened and is very interesting.

https://www.anuragk.com/blog/posts/Taalas.html


If that were the case, it would be reasonable to expect that companies like OpenAI or Anthropic, which are heavily indebted, would lose part of their business model, not because their models are bad, but because others will be cheaper and not as bad.


I think he means they will be commercially relevant and most AI compute won't be on GPUs.


Before they removed it, I was using groq Kimi K2 model for a chat bot in small community site/chat. It was really good, seemed to have incredibly vast general world knowledge and the fast speed (400tok/s if I remember right) meant that chat users got a response instantly which was a much better experience compared to other SOTA models at the time.

On the bright side it looks like Cerebras might be serving Kimi K2.6 at 1000tok/s soon https://www.cerebras.ai/blog/cerebras-kimi-k2-Enterprise


Those were amazing times. You could vibe code an entire prototype in seconds (200 tps). With Qwen3.6-35B-A3B and MTP, you can program at that speed on a single GPU at home now, but Kimi K2 is of course much smarter at almost 30 times the size.

I'm also looking forward for the Cerebras Kimi K2.6 release, which should be even better at 1000 tps. It is hard to overstate how important speed is for programming. Instead of having to wait for a few minutes until a task is done, it is just done instantly, and you don't have to context switch from whatever else you were working on while waiting.

I hope they will make it available to regular customers.


But too much of a speed doesn’t allow you to build up the context as the llm is working, it’s a two-edged sword.


Cerebras are only serving kimi for dedicated endpoint customers; for that you need a >$5m annual deal with them

Cerebras also seems to be killing off their regular APIs, they're deprecating models and GLM is still stuck on GLM 4.7, a whole 2 versions behind.


I was quite baffled they removed it and didn't double down on Kimi and serving the latest models instead.

Thanks for the tip, looks fire.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: