>Understanding what is going on with AI productivity is … frustrating to say the least.
Agreed. I think one of the hardest things about it is that productivity != value. You can push all the code you want, but if it's not driving revenue up or cost down, it doesn't matter economically.
Here is the best data I've been able to find. An observational study of 4000 teams over 2 years across many different organizations. Data gathered from their task management, version control, and CI/CD tooling. Critically - this is not survey data. It's much more direct measurement.
Are we plotting against cost? How is the capability advancement vs dollars paid for development?
By my read of the (very sparse) data, we're getting linear improvements in capability for super-linear increases in costs. [1] Indicates that by 2027 models will cost $1 billon to train. Dario estimates that model runs will cost $10 billion in 2026 [2]. That to me indicates costs are potentially growing faster than capability. Maybe by quite a bit.
If the value prop of LLMs doesn't prove out, that won't last. I'm of the opinion there is no data that shows actual economic value being delivered by models. The best data shows that LLM use might be destroying value [3].
I appreciate the data here but I don't think the read is quite right;
Saying we have linear capability for super-linear cost compares an unbounded variable (dollars) to bounded instruments (because benchmarks saturate). On unbounded measures, growth is exponential; you can see METR time horizons double every ~4-7 months (https://metr.org/blog/2026-1-29-time-horizon-1-1/). And capability being proportional to log(compute) is what the scaling law predicts.
Epoch puts training cost growth at ~2.4x/year as your link shows. Meanwhile cost for fixed capability falls ~10-40x/year (https://epoch.ai/data-insights/llm-inference-price-trends), and lab revenue is growing ~10x/year! Anthropic went from $1B to $9B to $30B+ run rate in ~15 months, OpenAI ~$25B.
On [3]: the "destroying value" conclusion flips sign on an assumed 15% baseline rework rate. The report's most direct metric is +16% merged PRs per dev. The RCT evidence is genuinely mixed (METR: -19%, with n = 20 and Claude 3.x; Cui et al: +26%) but its just super hard to do this well, I think Faros stuff was pretty cool, I haven't seen this before so thank you for the reference.
Maybe. There was a great comment in the thread on Fable 5 yesterday about benchmark comparisons between Fable and the latest opus models. here it is: https://news.ycombinator.com/item?id=48464600.
You could be right, but this is the most direct benchmark comparison I could find and it's not that strong.
>the "destroying value" conclusion flips sign on an assumed 15% baseline rework rate. The report's most direct metric is +16% merged PRs per dev.
I discuss this directly in my analysis. There's also an 860% code churn increase ratio. You only need 9% of that to be allocated to wasteful rework to drive throughput flat to the 15% rework baseline. Not to an assumed ideal state where there was no rework.
But even if it were not true, a 16% throughput improvement is pretty weak given the investment - especially given the direct evidence of quality degradation. IMO.
I appreciate you reading my stuff and taking the data seriously. Thank you.
> But even if it were not true, a 16% throughput improvement is pretty weak given the investment - especially given the direct evidence of quality degradation. IMO.
n=1 but at $JOB we have throughput quotas now, and what is happening is that teams are just finding lots of busywork (renaming things, gardening of ai .md files, rewriting uis etc) and also dividing prs into smaller chunks to match the quotas... so even "throughout increase" doesn't say much if its not for improving the customer outcome (ime anyways)
Yes I've seen this before, and while the critiques are fair and high quality (and unfortunately not unique to METR) we're missing the forest for the trees here.
First of all, if you take the articles critiques and work out the implications on the METR graph, all you're doing is shifting the curve up or down, it doesn't change the fact that progress is scaling exponentially. While it is technically possible the universe could be throwing a massive pathological curveball to change the conclusion from METR data (which is we've been seeing exponential growth over the last 6 years), I think that seems very far from likely. The fact that we see the same behavior from a variety of sources over a wide variety of tasks and domains is a pretty clear indication that METR while certainly far from perfect is actually painting a consistent picture at least in terms of the rate of progress.
You can look at ECI for a summary benchmark statistic, which does NOT use METR's benchmark, and you see a similar trend. Same with SWE-bench where the task distribution is far more in domain for real world problems. It is a bummer that this METR data can't be better funded. It would probably take $1M or so to really beef it up properly which any of these labs probably have in their couch cushions.
>By my read of the (very sparse) data, we're getting linear improvements in capability for super-linear increases in costs. [1] Indicates that by 2027 models will cost $1 billon to train. Dario estimates that model runs will cost $10 billion in 2026 [2]. That to me indicates costs are potentially growing faster than capability. Maybe by quite a bit.
This is true and well established.
As long as you get any improvement whatsoever, it is worth spending to train since it pays off during.
Imagine training was not $1 billion but $100 billion but the performance improved by just 10%. This is still worth it because you can squeeze out the profits across years and years right? The improvement is ever lasting.
> The best data shows that LLM use might be destroying value [3].
This is basically a conspiracy theory and if you really believed this, you should not have led with "How is the capability advancement vs dollars paid for development?" because if there were no value, it doesn't really matter how much you invest.
I think this is pretty uncharitable, especially when I've provided you with a dataset you can evaluate yourself and an argument you can review for logical inconsistency.
I have worked quite hard to locate data that supports your thesis, I can't find it. I've at least gone to the effort of documenting that search. Before you throw around such strong convictions, I suggest you actually look for yourself.
But what’s interesting is that you are commenting on a post where Dario is suggesting that LLMs are so extremely powerful that they can take over, help synthesise bioweapons, help in warfare, help in drug discovery — the whole post here is to try and regulate this. If you believe AI can’t even create positive value let alone discover new things then your problem is somewhere else and not in something like “but training costs a lot”.
So it is absolutely strange and contrasting to see you believe that LLMs are so weak as to create negative value while the CEO is asking about regulations because AI is too powerful.
I don’t think I can convince you that AI is actually that powerful.
But let me ask you something directly: if you believe what you believe, you should also acknowledge that AI doesn’t need regulations in the context Dario is proposing since obviously AI can’t do anything he predicts. Do you agree?
> So it is absolutely strange and contrasting to see you believe that LLMs are so weak as to create negative value while the CEO is asking about regulations because AI is too powerful.
You wouldn't ask a chemistry professor to write code. So just because LLMs create negative value for software development doesn't mean that they can't be helpful for bioweapons synthesis, especially considering the range of chemistry and biology sources Anthropic would have fed to its LLM that wouldn't be publicly accessible. The LLM doesn't even need to be particularly accurate so long as the amateur bioweapons researcher takes adequate precautions before following its instructions and does some background research beforehand.
This is a ridiculous stance to take. That LLMs are simultaneously negative value but can also help synthesise bioweapons. It’s the sort of stance you take when you already feel ideologically against AI. I don’t think it’s coherent.
It's more about information availability rather than intelligence. An LLM has had access to more information during its training period than you'd ever even come across over a hundred lifetimes. It has been trained on billions of books and articles across every single subject that exists on the planet. Can you imagine what real intelligence could do with all that information?
I'd love to understand why. This would be valuable feedback for me as I try to make my writing and exposition better. Also, if you have other data, that also would be valuable for me to know.
>if you believe what you believe, you should also acknowledge that AI doesn’t need regulations in the context Dario is proposing since obviously AI can’t do anything he predicts. Do you agree?
I think you misunderstand my beliefs. On net I think how we're using LLMs destroys value. That doesn't mean no one ever gets value from LLM use.
My particular point about trillion dollars is - the main place Anthropic, OpenAI, and - hilariously - SpaceX think they will drive value creation is in enterprise applications. In that domain I think the evidence is very convincingly negative. I'm certainly not the only person who thinks this. It's pretty well accepted in economics right now that there is no observed organizational level productivity improvement. Lines break down on whether it will show up eventually or whether we will wait forever.
My belief about LLM value is that it's most useful for individuals and small teams. Places where coordination and trust are easily established and feedback loops to value creation are tight. They are "short range" as it were.
Their value starts to erode as soon as a user becomes disconnected from the point of direct value creation. Which is pretty much everyone who works inside of a large organization. It becomes negative at pretty small scale, IMO. I do think there are patterns of use that could drive value at these scales. I talk about that in my post.
On Bioweapons in particular, I could see small teams of people working to build something very dangerous. Having spent my formative academic years in a biochemistry and microbiology lab though, I do think the danger is overstated. Papers are not know-how or equipment. There's a lot of tacit knowledge that can't get written down that is super hard to acquire.
But, I'd be happy for us to regulate AI for dangerous applications.
My question would be - why would Anthropic build something they so clearly think is dangerous? If they were really building something deserving of the valuation they have, why build applications like this?
To my eyes - it's super weird that a company would build something they think is dangerous and turn around and beg the governments of the world to stop them. That's really strange behavior from my perspective.
I went through your post in substack (I think that's what you were referring to).
> I'd love to understand why. This would be valuable feedback for me as I try to make my writing and exposition better. Also, if you have other data, that also would be valuable for me to know.
I think it comes down to few things
- you took a single report that agreed with your statistics, for the sake or argument lets say I buy it completely
- you suggest that net value is lost simply because there are more incidents. this is a big jump
- you say that historically different technological improvements may have had similar patterns but this specific one is different because AI is stochastic
So it all really rests on you finding one distinction with AI and then disagreeing with the past trends.
I agree AI is stochastic and I'll put it this way: it is a high variance bet but it pays off. This is a bit hard for people to understand -- its a tool that works sometimes really nicely and fails other times. Overall you are better off using it but you need to use it enough to reduce variance.
Let me ask this: if you are so sure this won't lead to enterprise level productivity, how do you think this will show in macro trends? Surely you must believe that the valuations must drop wouldn't you? Can you come up with a concrete future scenario that would vindicate your opinion that AI doesn't make enterprises more productive?
> My question would be - why would Anthropic build something they so clearly think is dangerous? If they were really building something deserving of the valuation they have, why build applications like this?
I think this is fair and interesting question. Here is what I think they think: If they don't build it, someone else might do it. And they think they are more moral than others. If they have a head start they can set the political and regulatory landscape.
>you took a single report that agreed with your statistics
These are not my statistics. I'm not affiliated with Faros at all. I built an analysis on top of their reporting.
And, it's also not one report. DORA has tracked statistics with respect to throughput and quality as well. Those indicators are flat for throughput and negative for quality. The throughput flatness is also supported by the shovelware data.
>you suggest that net value is lost simply because there are more incidents. this is a big jump
I don't think it's a big jump at all. Incidents and bugs drive rework. Rework has to be subtracted from throughput. Product throughput is the only thing people pay for.
This type of analysis is done all the time in manufacturing and devops. Here's a link for you: https://reworkcost.com/benchmarks. I'm not bringing novel intellectual ideas to the table here.
Faros reports a 16% throughput improvement on PRs. They also report an 860% code churn increase. If you assign only 9% of that increase to wasteful rework, then the absolute throughput improvement disappears. This is a very simple, straightforward analysis of the operations data reported by Faros.
> - you say that historically different technological improvements may have had similar patterns but this specific one is different because AI is stochastic
I'm saying LLMs are unreliable. I think we agree on that front, you say:
>I agree AI is stochastic and I'll put it this way: it is a high variance bet but it pays off.
What I'm disputing is the "pays off" statement. That statement is amenable to validation with data. In my view, the data is saying it doesn't pay off. I think it says that very clearly. Across distinct lines of evidence.
>if you are so sure this won't lead to enterprise level productivity, how do you think this will show in macro trends? Surely you must believe that the valuations must drop wouldn't you? Can you come up with a concrete future scenario that would vindicate your opinion that AI doesn't make enterprises more productive?
I think LLMs can deliver value in the enterprise. I think the way to do that is to use them as quality checks and not as primary authors of intellectual work - like writing code.
Unfortunately, this use case would not support the expected 2-10x productivity increases that current valuations depend on. I do expect a major market correction in the near future. It would not surprise me if OpenAI or Anthropic are acquired. I think we're at risk of that happening within the next 1-7 months.
What would invalidate my beliefs?
1. Actual micro or macroeconomic data indicating economic productivity is increasing.
2. A Faros like observational study demonstrating sustained throughput improvement with significantly less rework and quality impacts.
I think I could be swayed against the market correction if the financials of OpenAI or Anthropic are strong. I'm anticipating they will be quite bad. I think Mythos was very expensive to train and I think the improvements in capability are sublinear. The inference costs are incredibly high.
I also have ideas about how Anthropic and OpenAI are trying to change their business models into enterprise transformation plays. Similar to Palantir. But this comment is already long.
>If they don't build it, someone else might do it.
No other players in the market other than US tech companies have the capital or the technology to train the models of the power of Fable. The way the Chinese model builders are building their models is by distilling from US models. So Anthropic, by building Mythos with all this bio data, has created the possibility that other actors can distill their models and do harm with them. (Not to say the Chinese are seeking to build weapons, but actors with their models might).
I'm happy to make that bet. Just not for money. I don't gamble at all anywhere in my life.
But I'm happy to write something publicly like "Simianwords was right about this prediction and I was wrong". Also happy for you to suggest alternatives as well.
Ok fair. In the event I’m wrong I can make a post in my anonymous blog that I was proved incorrect and I lost the bet. But I’m anonymous and a nobody so it doesn’t count for much but it’s all I can offer.
I’m not sure which bet we are talking about but let’s go for the stronger one and I’ll repeat it here:
7 months from now, combined valuation of OpenAI and Anthropic will be 10% higher than it is today inflation adjusted.
By “break” I don’t mean “won’t function”. I mean “won’t deliver on their value proposition”. A functioning product is a necessary, but not a sufficient condition for technology to have utility. I would defend vigorously that the generative value in LLMs is derived from their unreliability. Which is what the argument ultimately rests on.
Human brains can explain themselves. What's a black box about LLMs is that they can't. They don't know how they arrived. They didn't arrive at anything. They don't have access to their internal representations. They don't think.
Productivity is not value. It's quite possible for you to experience productivity improvements, and actual value to not be created. That is what I think the most robust data is showing.
Also, supposed productivity gains are dubious. I personally experience at best no productivity gains when using LLMs to write code, and sometimes it's an active drain on my productivity. There was that one study a year or so ago showing similar results. People are trying to say the productivity gains are there and undeniable, but that is not true. It is very much a subject of controversy whether AI helps productivity.
I can see an argument that the productivity gains are illusory / don’t translate to economic productivity. I’m not denying the possibility.
However, most of the engineers I respect have gone from being skeptics a year ago to convinced today. I don’t personally know any true holdouts any more. If there are studies that disprove productivity gains more than six months ago, I’m happy to believe that it was true of the AIs that were available at the time. But I’m going to need something much more recent before I disbelieve my lyin’ eyes where it pertains to the AIs available today.
There is an observational study that was published in March 2026 that followed 4000 teams over 2 years. It shows, in my view, exactly that the productivity gains don't translate into economic value.
If it was published in March 2026, even if the data was collected up to the day the study was published, 7/8ths of it would fail my “within the last six months” test. But I am looking forward to the results of future studies on this topic!
I get wanting to wait for more data. And thinking that LLMs have improved enough that this will change.
My view is that it's not really about how good the models are - it's about how we're using them. Understanding what you've built is an important part of value creation, and LLMs eliminate that.
Its funny, I've noticed the same thing, but did not come to the same conclusion.
I currently don't have work access to Claude Code, but most of my teammates do. Watching from the outside, the cycle seems to look like this:
1. Experience some success, which hooks you into relying on AI.
2. The AI keeps failing at some task, but you don't want to stop. Keep trying over and over again.
3. Run out of tokens and take a break.
Now, sometimes 1 doesn't happen. Sometimes 2 doesn't happen. 3 is a certainty though.
Now, if you told me that the productivity gain from 1 is enough to offset the loss from 2 and 3, I could believe you. But I also wouldn't be surprised if it didn't.
As I work with Claude more and gain a feel for its capabilities, I tend to run into 2 far less often, as I'll decompose my messages more for the current model limitations. The threshold also changes each release.
I’m going back to being a holdout, but it’s nuanced - My theory into why LLMs don’t lead to the colloquial definition of productivity would be something like - if code was never the bottleneck than generating code faster doesn’t result in more meaningful output.
Even if you take for granted that AI is as good as the best people say in writing code. And Ive spent a lot of time generating codes, I won’t disagree - Then the question becomes - does this change your daily incentives such that you reach for code as the solution to your problems rather than something else (coordinating with your colleagues? Product management? Planning and Design?
So from a holistic perspective, I think intentionally limiting your own AI usage is the best approach for maximum long-term productivity.
>So from a holistic perspective, I think intentionally limiting your own AI usage is the best approach for maximum long-term productivity.
I think this is right. They are much better applied as editors than authors, IMO.
The key thing is stay in control of your output. i.e. understand it thoroguhly. I think you let the LLM make decisions you don't really understand, you're increasing the likelihood of introducing defects that are expensive to address.
I’m not completely closed to your idea but if code was never the bottleneck why did so many organizations always feel so chronically low on coders? And of course this requires the AI to be no help at all with what is actually the bottleneck.
So, one is, it probably does depend on workflow. Some folks probably are doing things that can be accelerated by AI, and I think if you’re a small team with a good product head and know what you’re doing then AI probably helps a lot.
But what if the problem you’re trying to solve is the altogether too often problem of like getting teams that are dependent on you to upgrade the library they use. And what if the library is a breaking change, and last year they upgraded to the library on your advice and it broke production and now they’re suss and want to accept all changes, and integrating that library change isn’t in their critical path so they’re just not going to spend time on it, even if you submit the MR them. Even if you show them their tests pass after the change.
Importantly to the above, you probably need more devs to do more of the above in parallel. You don’t hire devs to write more code, you hire more devs to carry on the mental load of a broader scope of work. Even in the before times, so much code got stuck at the integration step.
But because all that is hard, instead you go and codegen to fix an obscure bug that sure makes a few customers happy, but no one thought was a limiting factor for paying your company more money.
It’s not that I don’t think AI can help, I think it’s a prerequisite for the job and everyone should use it. It’s more that I think in the grand scheme of things, people will bias towards using it for tasks that aren’t in the critical path - refactors, tech debt, bug smashing, tool building; and I think it could really help devex and that’s good.
But I think people are bad at knowing the difference between “my job feels a bit easier” or “I’m more productive” and “this task had an impact on the bottom line” and when you extrapolate that out to a whole engineering org, that’s where the productivity statistics get lost.
I’ll addd one data point to this is like this thread itself. So many people on AI skepticism threads point to their own subjective experience as evidence we’re not in a bubble, and sort of ignore the entire concept of economics. I’m not saying we’re in 100% in a bubble, but subjective experience isn’t great evidence of it.
And this is just sort of one of the factors, what about the increased cost and mental load of supporting more software? What about junior engineers who feel pressured to ship work but don’t actually learn the software engineering? What about lost context from not intimately understanding your software?
All good questions. I am not a big believer in claiming I know whether we are in a financial bubble or not. I just put it all in VT and we will see what happens. I know that the AI allows me to write code I couldn’t write before much more quickly than before but I admit that this may not help with organizational friction.
Although if this theory is true — that AI helps with coding but coding is not the friction point in organizations with multiple humans, even that should allow faster iteration by allowing one human to do more coding therefore reducing the size of teams required to make some programs. You should see good acceleration in solo shops too.
Yeah, again the answer is definitely not no AI. And I’ve seen the - you know - I can one shot whole applications that would have taken me weeks or months before. But it’s also much easier to one shot apps that aren’t in the critical path. I’m building all sorts of tools to make my job easier, but I still have to do the job.
I’m a platform engineer. The primary failure mode for platform engineers is building tools people don’t want. AI doesn’t really make that easier. Or it can but it can also make it harder by making it easier to chase down ideas that you don’t get traction on. And I think that - net balanced across the organization is probably why productivity gains get sort of averaged out.
For sure I think solo devs who have a system are seeing gains, as long as you can I think have the discipline to have a process that includes feedback and learning and your not just feeding off of dopamine hits of one shotting features but yeah. I mean for solo devs the code was never really the limiting factor, it was product-market fit and marketing.
So solo devs who have a system may be laughing themselves all the way to the bank, but we may not see a lot of net new solo devs.
But if code is cheap now then it’s sort of inherently devalued. 2 8 person startups can probably relatively easily find a dev with AI experience to rocket ship their code generation, which means the basic skills of talking to customers, change management, and building the right thing become even more valuable.
Even solo devs I wonder - almost every post-mortem of a failed company goes “I wish we had spent more time talking to customers and less time writing code”
Again if you can get the discipline right, maybe as a solo devs you can get more work done faster and spend more time with your family. That’s incredibly valuable!
But if you go and add a big new feature, or a second product - unless your community is primed for constant growth(no man’s sky is one community where more more more seems good) you’re just growing the surface area where all the other skills are more necessary.
From an economic perspective productivity is defined as the creation of value isn't it? Then if you "improve productivity" and does not create value in the end you're no improving productivity at all.
economists define productivity as gdp per hour worked. Like a lot of other economic measurements, its mostly a bogus number people use as an argument on why their politics are better than someone elses politics. You can have an efficient business located in a poor country making the same product and same quality as that same business in a rich country, the rich country will be more "productive" because local cost of goods is higher there (i.e. a restaurant in NYC is more "productive" than a restaurant in bangladesh).
Sure. But that's not, in my view, how most people use the word productivity when describing LLM use.
In my field - operations - productivity is usually described as some rate of production for a specific asset. 100 widgets / machine / hour - for example.
"My productivity is 3 PRs / day with the LLM as opposed to 1 PR per every three days". That's how I think people are thinking about it.
My point is that's not the same thing as value. I.e. what people will pay for.
You're correct, I just wanted to add that there is another definition that you may see used online, and it is very specific, and it's important to be aware it's NOT exactly the same thing most normal people mean when they say "productivity".
The report was not paywalled for me. It just required a work email. Which is totally fair from my perspective. Faros is providing a ton of value with the report. People do deserve to get paid - even if in collected emails.
You're right my analysis is at variance to what Faros.ai says. I think they interpret their data trying to rescue utility for the dominant patterns of LLM use.
But I think to anyone who is experienced with process improvement or queuing theory, their interpretation is clearly weak. Rework is a huge problem in queue systems, and they mostly just elide the throughput impact of an 860% increase in code churn coupled to a massive spike in bugs.
Obviously draw your own conclusions. But I don't think because I disagree with the interpretation of the people who originated the data makes me wrong.
That's possible, sure. But I think the answer is more likely in the numbers, not in just qualitatively saying AI isn't worth anything. Like if I pay $30k for an ounce of gold, I got value. Gold is worth something. But that amount of gold wasn't worth what I spent.
EDIT: In fact, parent comment has a link to some numbers.
[EDIT: Most] people don't want to go through the numbers. Ok. But there's a history here. When people don't want to see the numbers, certain kinds of things tend to happen.
I've posted numbers that indicate that productivity is becoming decoupled from value delivery. If you follow the link in my comment it reviews a pretty robust study of 4000 teams over 2 years. There is no product throughput increase.
Code acceleration is great, but.... something precedes that. Vision and strategy re. expansion of offerings and businesses. Once a firm reaches maturity in what it offers and is only touching the edges - this code acceleration is literally useless when you factor in all of the trade-offs.
This is a good thing - it means fat and slow incumbents are sitting ducks to be out-witted by creative and imaginative founders, which is healthy for a well-functioning economy.
Now the economics of existing frontier models are not sustainable - its looking like a mix of the airline (supersonic vs subsonic) and EV industry with China in the background providing decent offerings at much lower prices.
I admit that if a small team or an individual uses an LLM, it's likely they can create value faster.
I think as soon as you don't own the responsibility for the defects you generate with an LLM, their use starts to destroy value. Regardless of product maturity.
Yeah this part scares me a little. I imagine it scares everyone who is more than a couple of years out of school. I hear that "the solution to LLM tech debt is more LLM." That might be true, but it might not be.
I actually think this is precisely the reason LLMs can't be the basis for a technological revolution. Because it's only one way.
Like, if you have a compiler, and it has a bug. You can discover if that bug is influencing your code execution and patch it. You can go both up and down the stack.
With LLMs, there is no way to patch it's translation function. You have to rely on it to forward process.
I don't think there is any way to avoid us understanding our tech stacks.
If you are producing something that delivers a far better experience, irrespective of what's under the hood (see Claude Code et al), you will decimate an incumbent who is trying to use LLMs in the context of incrementally improving a mature product.
LLMs are suited for the development of revolutionary innovation, not incremental.
Agreed. I think one of the hardest things about it is that productivity != value. You can push all the code you want, but if it's not driving revenue up or cost down, it doesn't matter economically.
Here is the best data I've been able to find. An observational study of 4000 teams over 2 years across many different organizations. Data gathered from their task management, version control, and CI/CD tooling. Critically - this is not survey data. It's much more direct measurement.
https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways
Faros argues that teams are seeing about a 16% throughput improvement (PR merge rate) with heavy AI use.
I argue here that their data actually indicates negative absolute impact on throughput.
https://unessays.substack.com/p/talk-is-cheap