Hacker Newsnew | past | comments | ask | show | jobs | submit | yojo's commentslogin

If you can afford to eat the loss, don’t insure it.

Obviously insurance companies make money on average for every policy. If they didn’t, they would go bankrupt. So on average, every time you buy insurance, you lose.

So insure only what’s required by law, or things that would throw your life off track if they vanished.


Yep. Which is also why I find vision insurance ridiculous. No it doesn't cover me going blind, it covers a pair of glasses at best, more like 25% of one.

That’s nice in theory, but as these things get better and cheaper this kind of capability is going to drop from nation states to script kiddies. That future is coming, I don’t see any way around it.

We can round up all the bored teenagers we want, but it’s not putting the genie back. Better start adjusting our systems to account for it.


Ever read the Anarchist Cookbook? Anybody tech inclined with a hint of mischief in them, from a certain era, has. It's a list of all sorts of awful things you can do, mostly with household ingredients, and a few minutes. I think its overall impact on society was pretty much zero. Actually it may have been overall positive because I expect plenty of peoples first experience with things like thermite came from that book, and now there are all sorts of videos and neat experiments with such on sites like YouTube.

I think this is in part because most people, including awful, tend to be relatively morally inclined. But I also think because even with an LLM, doing things takes effort. And if you're willing to dedicate effort towards a task, there tend to be way more rewarding/gratifying things to do than try to hurt people. Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.


I suspect people didn't start blowing up stuff because they understood that would be bad, harmful, and also very illegal. Everything computer related somehow seems to feel less real or consequential to some people. And AI doesn't have this compunction at all unless we make really sure it does.


When I was a mischievous kid, me and all my mischievous friends had our stories of learning how serious fire and explosions were considered by authority figures. We learned fast not to do that or the consequences would be grand. These were usually small fires or firework involved pranks. So, yeah I agree with this.

Computer stuff has generally always been a slap on the wrist in comparison. Maybe it’s more punitive now. But also, it’s one of those things that maybe you get in trouble officially but at home and behind the scenes you’re friends and maybe your dad are laughing and giving you high fives. So young mischievous kids will totally go there because they’re not afraid of punishment if it is minor and it gives them a notch on their belt. If they can take down Amazon.com website for a day, we all know that’s a massive financial implication, but it’s also a faceless mega corp and quite tempting if you can get the bragging rights with only risk of a small punishment. (Note; I don’t know what the current crime/punishment for this would be, and whether it’s small is very subjective).

It’s similar to how some people gravitate or succumb to the opportunity of white collar crimes. Embezzling $10m from a company almost makes sense in a situation where that only gets you 5 years max prison. If you hide it well, you simply serve your time, and then retire in comfort. I can see how that makes more sense or is tempting to people than slogging through a lifetime of low income job as a bookkeeper just trying to find a way to save for retirement.

Most of these people would never consider robbing a bank. First of all, it’s not a $10m dollar opportunity. Usually not enough for anyone to retire on, or live more than a year or two really. Second, it’s usually considered a much more severe crime and sentencing can be very long, I’ve seen 30+ years. Third, it’s much more risky to your person. Getting shot and dying is absolutely possible.


IMHO it's about proximity. It's easier to be inhumane from a distance.


the anarchist cookbook actually results in a ton of things that are more likely to explode the user than any intended target.

TM 31-210 on the other hand will not: https://en.wikipedia.org/wiki/TM_31-210_Improvised_Munitions...


> Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.

This theory interests me.

I'd love to understand how different individuals within intelligence orgs have reasoned about the morality of their actions.


The FBI cyber teams will have agents too. It will be a glorious war.


That would make it the first one in history.


I wasn’t being serious


They still haven’t removed “old.reddit.com”. Replace “www” with “old” on any reddit URL and enjoy a relic from when the internet was less ruined.


The issue is that they just started forcing sign-ins on Old Reddit, allegedly because it's easier for bots to scrape. The ability to casually/anonymously peek at a post that answers your question is getting more rare.


Yes, but it now requires you to have an account.


And sometimes even to verify with Persona.


I was at Dropbox from 2016-2020. We were certainly trying to build a sustainable business, but there was a major identity crisis. Were we consumer web? Buy Mailbox and build Carousel, then shut them both down.

Maybe we’re Notion/Evernote? Buy Hackpad, plow a ton of money into Paper (which was legitimately good), then quietly deprioritize it.

Maybe we’re actually some kind of enterprise document productivity suite? Buy HelloSign. Plow a bunch of money into a desktop app. Pull more plugs.

A lot of smart people were trying. We made a lot of bets (too many?). None of them proved to be a second act, and the competitors eventually caught up.


I intend to write a piece going deeper into the failures; I would love to chat if you're open to it. Also, not to say Dropbox sucked or anything, it is just that the broader strategy and industry structure make it hard; if anything, the success of the first product made it difficult to evolve the business.


If you didn’t allow it, you wouldn’t be able to change models in the same conversation, as key parts of the context would be lost.

Wouldn’t surprise me if the providers just remove that ability and lock the model once the conversation starts.


Fair (I haven't been using the encrypted-reasoning systems, though this is common in open ones - I'm kinda surprised it's an option in encrypted ones too), though what they're doing here is cross-user replays in addition to cross-model.


100% guaranteed that this research just forced this to happen now.

Sucks.


It’s already patched according to the authors. Details were not specified.


or add some metadata and don't allow downgrading.


AFAIK no provider guarantees compatibility of reasoning traces, even in the same model generation, and we've in practice seen most of the big LLM APIs throw errors indicating incompatibility (at least transiently) when switching models. The only stable solution right now is to just throw away reasoning traces whenever a model is switched.


There seems to be an obvious choice to make here, should you give the users to decrypt and use the COT that they did not generate themselves?

This is only required if you want users to be able to share things with everyone and you are going for the simplest implementation.

If not you could try to keep a record of keys associated with a user, then when a new request comes in look through to see if the user has a valid key to decrypt the COT.

For explicit shares, just add the key used in that one conversation to the users valid keys. For global shares use the global keys. But that's adding more complexity to the system.


It’s about being able to change models mid-task. For example, I want to be able to plan using Fable but implement the plan using Sonnet, and that won’t work if this is implemented.


Or even my fable credits run out mid task and need to switch back to opus >.<


Prior to LLMs I never considered that I might have to make a resource-usage decision between hiring Star Trek's Data vs. his stupider brother B4...

https://memory-alpha.fandom.com/wiki/B-4


Star Trek is a post-scarcity society, those problems don't exist there unless you're in the middle of a crisis and on emergency power.

LLMs briefly seemed like this too, after subscriptions made the SOTA models too cheap to meter, but before they walked back on that and introduced quotas...


Actually, that brings up a good reason they can't fix it. Fable falls back to Opus when the topic is too "unsafe". That behavior requires traces than can move between models!


For plan it's relatively easy, just make the plan the artifact. The point is to ingest knowledge with one model and use it in another, and that is not necessarily easily expressible in natural language.


I suspect that there are companies with internal proxies that load-balance across keys, and they didn't want to break that when adding encrypted reasoning.


I really don't understand why server-side storage of the trace isn't a viable approach here, with only a unique key flowing to the client and back. Does it have something to do with how backend load-balancing works?


Makes no difference. There is a policy as to whether to allow use of a reasoning trace in a given context. Whether that trace originates from authenticated ciphertext or a backend database is basically irrelevant.


Good point, thanks.


Yes, this storage would be growing exponentially making the disk space and latency problems harder (add the disaster recovery/backups). I think the choice of using client side is not too bad if you ensure that its secured properly. Also the company can excuse itself from the liability of storing sensitive data on its servers, thats a big deal in itself to be compliant for enterprise audits

1. The down side is that it cannot be used across the clients even for the same user

2. Using the same encryption key was a bad choice here, a per user key would have solved this issue for sure.


Having thought about this a little more, it's clear that server-side storage is not compatible with Zero Data Retention (ZDR). However, in non-ZDR settings, it seems likely that the providers are capturing all that data anyway?

> a per user key would have solved this issue for sure

It would have helped with PII leakage, but not with plain-text trace extraction attacks, right?


Per user encryption key ties it with the user session (assuming you do authentication properly), no one else can access it. User being able to see the information is not really an attack vector in this case.

The compliance rules at times are outdated and people skirt around them by following the worded rule instead of the intent.


Encrypted state cookies solve real problems (server-side storage, latency, scaling) and are not the problem. The problem is insufficient binding of some of a session's encrypted state cookies and others -- insufficient binding of some session state to other session state. Here we have HTTP encrypted state cookies for identifying authenticate user IDs and maybe for identifying sessions / chats, while the reasoning traces are also encrypted state cookies but not HTTP cookies, and the latter are somehow not sufficiently bound to the former.

The fix is to either have per-user or per-session keys for encrypting reasoning traces, or write the user ID / account ID and maybe also session ID into the plaintext of the reasoning trace _then check that that matches the ones in the HTTP cookies when decrypting the traces_.


That’s incompatible with zero data retention and so you’ll lose a lot of enterprise customers.


Honestly, GPT 5.6 Luna is worth a look. It’s a reasonably good implementer at a small fraction of the cost. $100 buys a heck of a lot of it at API pricing.

Not sure about the OAi Pro plan, doesn’t look like the 80% Luna price slash made its way into the quota system.

You could also try tuning down the effort level on Opus. It makes a huge difference in token consumption and you might be able to get away with lower than you’ve set


At the beginning of the pandemic all the mills cut production anticipating economic collapse that didn’t come. Prices spiked on high demand and low production.

Mills did ramp back up, but it’s unclear to me if they used it as a chance to do so slowly/preserve margins. Lumber never got close to pre-pandemic levels.

Tariffs probably also play a role here. About a quarter of US lumber comes from Canada, and barbed wire is just steel with a little bit of processing.

For products with little value add there’s not anywhere for the tax to be absorbed, and no real way for domestic producers to quickly scale up, even if they wanted to.


Part of the problem during Covid in the US was US lumber mills are set up to do trees of a certain size. If the trees are harvested late they're too big, so we had to send a lot of lumber to Canada where it could be processed.


Claude Code injects a ton of tools into the system prompt, including their “memory system” that’s like 10k+ tokens. Depending on your task shape, this can easily double your task cost (e.g. a low context-using job that takes many turns, like a monitoring loop).

You should use --disallowed-tools to prune any tools not needed for the task. Note that this is also a perpetual game of whack-a-mole since they’re always adding new tools.


> Claude Code injects a ton of tools into the system prompt

It's so unfortunate they don't let you use the subscription with other harnesses anymore - since even if I used OpenCode they'd still get a bunch of useful data from the API calls, meanwhile I could stretch their tier limits way further.


The funny thing is that in some cases using a different harness with the subscription plan could actually be very good for Anthropic: e.g. if I were to use smol with Opus, it could use fewer tokens than CC for the same task. People on subscription plans burning fewer tokens is a good thing for Anthropic.

The only downside for Anthropic that I can see is that hitting your limits more often (while using CC) could make you want to upgrade plans, and a more efficient harness could keep you from doing that. But I can't imagine the cost (to Anthropic) of those inefficient tokens is worth it to them.


Subsidized Anthropic subscriptions seem to work fine on the oh-my-pi harness, somehow.


Calling per token usage of the US closed source labs has always been funny to me, we have a clear model of what it actually costs to host these models from open models.

Your subscription is not subsidised, it is just closer to the actual cost of the model…


Really? I immediately got a warning that the tokens would come from extra usage, and noped out immediately (so maybe pi lied to me?)


It's so unfortunate people don't realize it's cheaper to write their own well working agent instead of insisting with general purpose bloated ones like Claude.

200$/month is a lot of money on Luna/DS4 flash, like really a lot and the results are much better than clowning on bloated CC.

It's absurd how you have more and more organizations encoding their processes on LLMs and "engineers" (charlatan coders) don't even bother optimizing the tool they use most.


> the results are much better than clowning on bloated CC.

I won't argue with the cost effectiveness, but the results are very much not better. Opus and Fable are in a different league than DS4 Flash. Even GPT Terra, which I really like overall, sometimes gets stuck in weird loops and starts to do stupid stuff once its context window fills up. Whereas I can more or less trust the big models to just Do The Thing™ on the first try.

With that said, you get way more value out of a GPT subscription than you do from Claude, partly because of the ability to use more efficient harnesses.


You're confusing models with the agents.

You can use opus or fable if you please, I'm advocating for writing your own agent instead of used generic bloated ones like CC.


I'm not confusing anything since you can't use custom harnesses with your Claude subscription—you have to use Claude Code. So as far as the Claude sub is concerned, the models and the harness are coupled.


Thanks, you make me feel better for spending the (fun!) time writing my first harness in Emacs Lisp with a ‘Emacs UI’ and later writing a command line harness in Common Lisp. Am I more productive with my own harnesses? Not yet, but the second harness I wrote in Common Lisp is promising, and I might write a little book just on this project to encourage people to hack what I wrote and make it their own.


I have yet to spend $200 on DS4 after two months of using it with an entire team.


It's cheaper still, and just as effective, to write code yourself instead of having the LLM do it for you.


According to ccusage I use about $5k of tokens on my $100/mo Claude Max plan and only hit 5-hr windows where I have to switch tools for a couple hours about once a week


The system prompt part is surely cached across all users.


You still pay cache token costs on API calls. Cache cost/token are 90% lower, but you pay it every single turn.

I’m not sure if they let you skip the cache write cost on the first turn. That would imply cross-user caching infrastructure or special casing the default system prompt to give you a discount. Maybe? Away from the computer but you could try a “hello” in a fresh session and see what was billed.


Its inexpensive to reserve KV cache for the first turn and would benefit users, given the first turn already requires costly locating and allocating a model slot for a user.

So yeah, when they banned tgird party harnesses there was a technical and $ case to have.


For Anthropic models, yes. But OP was using Claude Code with GPT.


the system prompt still takes up useful space in the context window and steers the model into unnecessary actions and over-thinking patterns


ty re --disallowed-tools

for coding agents 'shell' is often all you need (just make sure the environment has the necessary tools)


appreciate the tip, kind netizen


I read the balls as “houses”, though the phallic spire emerging from them dead-center is also very on-brand


I must admit, I have a natural tendency to overlook C&Bs, but solid catch -- it's there. Perhaps that explains the Mona Lisa redaction; Grok probably put more effort into that one.


I have been working all day every day in Claude. I loathe their bug-ridden UI. Every release is a new crop of bugs, sometimes the old ones get fixed, usually not.

Any kind of scrolling back, copying text, using their menu system - basically anything that isn’t typing characters has had/still has unaddressed bugs.

OpenAI shipped a competitive model and I’m over in Codex now. I have yet to hit a bug.

If you’re holding the SOTA crown, people will put up with your buggy mess. As soon as that crown slips your pile of trash becomes a huge liability.


> basically anything that isn’t typing characters has had/still has unaddressed bugs.

Oh, that's buggy too. I just tried Claude Code on win10 powershell, and the first typed character goes in the wrong spot and can't be backspaced.

It is by far the the least reliable program on my machine, and every time I have to interact with it I feel like walking in eggshells.


> I loathe their bug-ridden UI.

So weird that the same exact people telling you that programming careers are now obsolete are the same group who haven't been able to fix screen flickering bugs for like a year...


Oh, and the memory use! I run a lot of concurrent sessions. 3 gigs for a terminal window is ludicrous.


The vim mode in that text box was a mess for sure


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: