Hacker Newsnew | past | comments | ask | show | jobs | submit | strulovich's commentslogin

I recently needed to buy some hardware for a piece of furniture.

Ran Codex, it found it for 18% less than what I found in the top Google results. It did it by finding smaller shops, applying a discount code, subscribing to a newsletter for a better code after approval, and took into account the shipping (by placing it in the cart and going to checkout) all to get me the best price.

I’m guessing without it I would have spent much more time on it and paid the original price I saw.

If you use AI agents well, they can easily save you more money than they cost, and saving money is something most people are pretty excited about.

(Disclosure: OpenAI employee)


>to get me the best price

How do you know it's the best price ?


Found the evals fan!

  Hi Omer,

  I think that is a use case where I'm sure an AI agent does well
  right now, but I have 2 follow-up thoughts based on it and what
  it is means that this is one of the best use case for AI online
  commerce so far.

  My first thought was that this is a quite limited application as
  most of shopping (online & retail) is not spent comparing prices
  and is instead spent on finding what product to buy in the first
  place. I believe this "shopping fun" and exploration of multiple
  options is not going away any time soon as those are some of the
  most attractive parts of the whole experience.

  I also think that the price comparison websites that exist, e.g.
  geizhals.de here in Germany (includes most e-commerce shops, and
  shipping information), are pretty good already and AI only being
  a better version of that would be quite a sad turn of events and
  with ads (and commission systems?) potentially coming to ChatGPT
  in the near future the incentives to find the best price are rly
  misaligned and I have doubts that people will trust results from
  ChatGPT. We already saw this as Google's sponsored links evolved
  to show whoever spent the most for the placement. Paid providers
  without ads and commissions might fare better though.
edit: paragraph 3 is already happening: https://news.ycombinator.com/item?id=49563386

I mean, that's cool and all, but the numbers are really going to shift when it accidentally goes off and orders that same hardware from every vendor in your local region and the top 5 online results for comparison.

It's the same problem as all other LLM solutions (that I hope OpenAI is working on!) it's non-deterministic, and there's no way for the user (or model provider) to know what the distribution of possible outcomes is. This just gets compounded when multi-call harnesses come onto play.


Thanks. This is genuinely a cool usage example.

How did you run this? Web interface, desktop app, CLI?

How did you complete the final transaction?


its going to get worse though as cloudflare keeps blocking more and more

I‘m using mostly the browser automations for things like that now. Same for research - ChatGPT is banned from reading many pages, but Codex can read anything I can.

But isn’t it funny that Cloudflare is blocking AI on their pages, but on the other hand is researching and marketing things like „you can put a browser in a CF worker“


Their browser workers won’t get blocked. Same with all big vendors, keep out small competitors enjoy access yourself and sell it to a select few partners.

The sites that do that won't be getting money from my and others' agents. Guessing that's going to become more and more of a problem for those sites.

I expect a small site that undercuts the top Google results by 18% with a sign-up discount probably isn't profitable on those orders, so blocking agents would save them money - it's not like someone using agents in that way is going to have any loyalty to shopping from that site in the future.

If no one can find your site because you block agents, that won’t do much good either.

Yeah, I mean if that's the case, then said small sites are probably headed for extinction, and agents won't have anything to work with other than whatever the top Google result is anyway.

And yours and other's agents will probably remain an insignificant and invisible customer base anyway.

My crystal ball is as good as anyone's, but if "agentic shopping" ever becomes mainstream, you can be sure that the vast majority will ask their phone (i.e. Google, i.e. Google Shopping) what the best price is anyways.


Sure, Google could be their agent, or ChatGPT, or Claude, or whatever local model the person is running. The big guys might all have the results cached so they don't have to rerun the crawl. Whatever their choice of agent, it seems pretty clear that almost everyone's going to use them, they're way too useful not to.

The best trick I have after asking it nicely in all sort of ways is:

1. Have it build a scoring script that penalizes words outside a simple English list and approved jargon. Penalize sentences over 15 words as well. Add whatever else.

2. Run it in a loop to reduce the score while preserving intention

This works much better than other ways I’ve tried. Of course it costs more. And I would apply it only to the output to the user, not the thinking process (I think the AI thinks better with their crazy English)

Of course, sometimes nuance is lost by this process. That’s just the nature of making things simpler.


I haven't tried this with a score but I have a simple skill with some examples of PR description changes and good PR descriptions I'd previously wrote and I just run it on the description.

It does cost more but I haven't tried cheaper models to see if they can get the same results. Curious if anyone else has.


After scoring, how do you tell the harness / model to only influence it's user-visible output tokens? Is there a deterministic way to specify this or is it a plain-text instruction in a hook or skill?


Aren’t Waymo prices above Uber’s? That’s what I’ve seen in SF.

That is to say that I don’t think the economics work. You’re better off offering free Ubers. And I doubt that will solve many issues.


Waymo charges more than Uber to supply constraints, not bad unit economics. People (like me) are willing to pay extra for the product Waymo offers, and Waymo is itself supply constrained and has no direct competitors offering driverless rides, so they get to do a bit of monopoly pricing.


    > Waymo charges more than Uber to supply constraints
To understand this better, is Waymo limited in each location by the number of cars they are allowed to drive? Or, is Waymo too slow to expand the number of cars on their side? As an outsider, if you have very good demand for your product, why not expand? I must be missing something obvious. (Again, I don't write this to doubt your explanation.)


AIUI the issue is just scaling pain. No one has ever run a global network of first-party-owned taxis before, and they're learning by doing, at the same time as they're also solving problems in ML/robotics to make the robotaxis work more reliably in more contexts, and also figuring out how to manufacture reliable robotaxis at scale.


There is a limit on how many they are allowed to operate.


I tried to research this limit. Here is what I found about San Francisco:

    > Waymo is not subject to a regulatory cap or hard limit on the number of cars it can operate in San Francisco.
To be clear, the regulator in question here is California Public Utilities Commission (CPUC), which is a state-level regulator. Apparently, SF local officials would like to limit the number of vehicles but it not allowed by division of powers.


What's the product vs an Uber? Genuine question.


Drives better, isn't distracted, doesn't get lost, care are always clean, will always take you where you want to go day or night, and yeah, no human.


Yea, it not that most uber/lyft drivers are bad (most are good), it’s just that all Waymos provide a good experience.

It’s the same reason many people will prefer a chain restaurant to an unknown mom & pop. By knowing what you’re getting, you’re willing to give up the potential for something nominally better to avoid something much worse.


    > car[s] are always clean
This is a real surprise to me. I don't write that to doubt you. How do you think it works in practice? My point: If you have a human in the car, they can keep a closer eye on the car between rides and keep it clean. Without a driver, you lose that. How does Waymo know when a car needs to be cleaned?


Waymo regularly brings the cars in to central service depots for cleaning and battery charging. Also, as a passenger you can report that a car is dirty and it will be pulled out of service for immediate cleaning (and you get a small credit towards your next ride).


Waymo has done a great job of making its rollouts into new cities smooth. I have to wonder what will happen once they pivot to profitability in a few years. Will they be willing to lose revenue pulling so many cars out of service? Will they optimize the human costs by not having so many humans on staff?

Maybe they'll kill Uber first and then drop the quality.


"Doesn't get lost" is generous. they always cause so much caos in SF


I don't have to listen to the driver argue with his girlfriend across the entire span of the bay bridge, I don't have to worry the uber reeks of week (or worse, cigarette) smell on the way to an interview etc etc. Given the choice I will choose driverless every time.


Do you report these complaints using the Uber app? To be clear: I have no experience riding with Uber.


I have written about this before in similar discussions. I think men just cannot comprehend how much women are willing to pay to guarantee that they can avoid a scary situation in a car with a stranger driver. Ask any woman: Have you ever had a scary experince in a car with a stranger driver? Probably 100% will say yes. I would say Waymo's primary customer will be middle to high income women in urban environments.


Also middle/high income parents. Shuttling kids to and from school and activities is the real killer use case for driverless cars. They're actively working on getting this legalized.


No human in the car?


This is only a product for anti-social nerds. Most people will not pay a premium for this (this isn't waymo's product).

Edit: lots of people commenting on here how they're normie and they would still prefer waymo. Read what I wrote again - I didn't say normies wouldn't prefer, I said normies wouldn't pay a premium. So unless waymo is planning on being a luxury car service (which it's not) that can't possibly be their intended product/market.


They absolutely will. I'm a very social person but when I'm going home from a date with my girlfriend I always want a Waymo just because I don't want a 3rd person involved. And sometimes I'm having a bad day and don't want to have to make small talk. And it's nice that I can talk on the phone without disturbing someone.


Safety is normally ok in an uber but in a Waymo it’s much higher. Whether I’m talking about driving safety or personal safety from the/for the driver, you pick. But generally you can charge more for a safer product.


This comes up in conversation at office events, every woman in our office always agrees they will choose waymo over uber every time, due to the lack of (presumably male) driver.


My extremely social not-tech nerd wife that hardly uses FSD in her car got HOOKED on Waymo after the first ride. Same with my parents.


huh, definitely will pay more for not having a driver to be constantly on the phone.


yes yes you are the product.


Waymo is more until you factor in the tip for Uber. Ive also found that waymo is a lot more correct in estimating time to pickup in SF. Because Uber/lyft drivers can turn you down (presumably) I find that "your Uber can be here in 4 minutes" when I open the app turns into "but it will actually be 20"


I think Uber just lies about estimated pickup times and Waymo doesn't. Partly because I've never had an Uber arrive faster than their estimated time, whereas Waymo arrives early pretty routinely. And also several times now I've had Waymo say "it'll be 20 minutes" so I book Uber because it claims a much shorter wait time, and then Uber winds up being about Waymo's estimate anyway.


Who tips on Uber? One price up front and paid, that is the deal. Anyone tipping on something like that is their own fault.


I dunno. Hard to say "I tip my Uber driver because I understand how little the drivers get paid" if I'm also moving to waymo instead. Hard to find ethical consumption under capitalism.

But since you asked in kind of a flippant way: who tips on Uber? People who aren't sociopaths, I'd say.


>who tips on Uber? People who aren't sociopaths, I'd say.

Non-americans don't. To me, being socially expected to tip for a Taxi feels like a parody from a sketch comedy. ditto for tipping barbers or expecting tips over 10%. The only time I've tipped a Uber was when me and my friends were at risk of losing a train. The driver (unprompted) went pedal to the metal to make it on time, so afterwards I gave him a couple of bucks.

Tip culture is bad for rideshare drivers in aggregate anyways. The apps carefully manage the pay of drivers on each market so necessary supply is guaranteed at minimal cost. If a expectation of tipping exists, the platforms simply pay less. Now drivers also have maximize for tips instead of just driving the goddamn car.


Fair enough- outside of the US, workers may get paid fairly. In the US, many service workers are in a system where the employer has externalized the cost to pay those employees to the customers. The system is awful, but I'm also not going to not tip if I can afford it.

FWIW, the real problem isn't tip culture, it is the power imbalance between labor and the company that is using that labor (without employing drivers), but I dunno how much of an audience that kind of argument has on this particular site.


> outside of the US, workers may get paid fairly

I don't think that's why there is little tip culture in eg east Asia. It's just... the price the business expects you to pay is mostly advertised upfront, it's known to everyone involved, why would you pay more if you got exactly the service you have ordered? It simply does not make sense.


My country has stronger worker protections (for now) than the US, but employers would be happy (and do try!) to externalize pay to customers via tip. It doesn't work because people find it insane and refuse the US inspired nag tip screens.

The US will never fix the externalization of salaries until the public refuses to tip at ridiculous percentages and situations. You'll see a 40% tip expectation at restaurants, your local dentist will install a tip screen. Just tip 10% at sitdown places and refuse emotional blackmail. If you feel any guilt, just make the habit to tally the dollar amount of "normal" tips and make monthly donations to your local food bank.

EDIT: My point is that tipping culture is unrelated to worker rights. Many countries with large informal economies (i.e minimum wage isn't enforced for most people) have no expectation of tipping. The reason Norwegians don't tip drivers isn't because they have a wonderful safety net, it's because there's isn't a social pressure to do so.

Mixing charity and payment of services is stupid and a moral hazard. Servers have shit wages in the US by law because people tip and the same is true (via algorithm) with rideshare drivers.


I'm not a sociopath but I believe the concept of tipping to be terrible. Everything should be in the price. up front. what's advertised should be what leaves your account. If you get bad service, don't use them again or leave a bad review. Get good service then thank them and leave a good review. The rest of the world works this way and it is simpler and less stressful.


Not tipping being "sociopathic" as if you are even actually using that word correctly must be the most American thing I have ever read. Most of the world does not tip like this.


This is not a very charitable explanation, it takes politicians at their word during a war. (One should not do that, and you can refer to Putin’s language at 2022 as a parallel example to Trump’s)

The initial US goals clearly were: 1. Regime change 2. Denial is of nuclear weapons

It’s also clear these goals were not achieved. So the US changed tactics and goals. (Same as Russia no longer plans on capturing Kiev it seems)

Most likely the US is stalling for time due to oil markets and has the same intentions as before, limited only by current capability.


I think regime change is likely to happen within two years. Just not in Iran.


According to The Economist, the Iranian theocracy is no longer in power, the IRGC is. Still, not the regime change Trump was hoping for, that's for sure.

> Khamenei’s killing has accelerated Iran’s transition from a theocracy to an ambitious nationalistic state dominated by military men. The irgc appears to wield power with few constraints. Clerics who challenged its influence—including former Presidents Mohammad Khatami and Hassan Rouhani—were conspicuously absent from the [funeral] processions.

https://www.economist.com/interactive/middle-east-and-africa...


>IRGC

Forgive me its been a while since I read about Irans political structure, but my understanding is that the IRGC is supposed to take over in times of succession crisis, and sort of take any measures to guarantee the islamic revolution.

The test is supposed to be that they hand back power sometime after the crisis.

If you assume Khamenei Jr is still unwell, and there's still a spot of bother regarding what his succession would look like, and the civilian government is still a bit in shambles, the IRGC taking over seems very easy for them to justify. Whether they hand that power back willingly is another matter that remains to be seen.

The problem here is that Trump bombing Iran is going to keep them in power longer. The IRGC being in charge is going to keep Trump bombing them. I dont see a way out of that spiral on either side.


> Trump bombing Iran is going to keep them in power longer. The IRGC being in charge is going to keep Trump bombing them. I dont see a way out of that spiral on either side.

Your statement actually makes the way out of the spiral very explicit: Trump must stop being in power.


Where I live (NYC) putting altered images like that has been the norm for more than a decade.

It’s just used to be more expensive to hire someone to do it for you.

The altered images always e free stirs the same bright walls and grey magazine style furniture.

AI is just making it cheaper, but this was bound to happen.

(Images altered this way do have a small watermark stating so)


Also NYC. A classic was mounting a bright light outside a window so it appears as “sun-drenched” as the description claimed.

(Unrelated, my favorite one was getting to the apartment and learning the “bedroom” was a flex wall in the kitchen)


AI has very uniquely made creating these faked/impossible layout images one of the cheapest & easiest things you can do at the moment, even if it didn't introduce the concept. Simultaneously, AI has had very little cost reduction impact on much else. This change in relative balance is how AI has created the new version of the problem and it's not apparent how this imbalance was always bound to occur without AI.


  > e free stirs
Features? Which TTS are you using? I was until recently using Gboard but it's been getting unusable lately.


The Linux kernel is not in any way at top of big projects. A kernel, as the name suggests, deals with specific issues and tries to remain small.

The world’s biggest software is usually built over endless adapters of different data and a need to reconcile endless edge cases with laws, regulations and real world complexities.


Now let’s see you do this with 40B. The f you can do that I will be intrigued, since it sounds like you solved a bunch of complicated economic problems.


I could sell someone a hundred billion dollars for 40 billion dollars and have 40b in revenue. It would never make me any money.


Of course you can sell $1 for 90¢, but your unit economics look terrible. If you want to seriously critique Anthropic you need to explain why their unit economics are bad.

Given that they just filed a (confidential) S-1, we will get an answer soon enough.

I expect it's going to look a bit more like selling 38 billion dollars for 40 billion in revenue^ than your example.

^ Some other caveats about how they're marking their p&l, but I think if the growth continues and they have a durable moat^^ then this will look like Amazon and be able to pivot into higher margin stuff

^^ haha this is the biggest condition, but I'm optimistic


The truth is that the backlash is stronger against E2EE.

Avoiding E2EE is less PR hassle to Meta than having it.

As a supporter, I must admit the E2EE interest groups have not had the upper hand recently (if ever)


They are blocked from insider trading

I think this in the previous comment undervalue just how many more complicated ways there are for corruption to bubble.

Assume super strict rules. Consider this example to circumvent them: Senator-elect A makes an LLC and invest their money (with some friends) in it (before taking office)

They can’t even talk to the money manager. But the manager can see their actions, and the senator can know their ow investments and work in their favor.


UnusualWhales begs to differ


Because research on this topic supports it. Happiness and wealth are correlated.


Only up to a certain point, no? I remember it was something around 100k USD, maybe 10ish years ago.

This is pretty intuitive. Its nice not to have to worry about money, but what is the difference between having 1M NW and 100M? If you're a mentally normal person, it just more mental burden.


Recent research disproves the old limit which has grabbed headlines like that old half a glass of red wine is good for you paper.

And also. Up to a certain point is still a correlation. Getting a lot of downvotes by people not knowing what a correlation is.


Really? Last I read the correlation breaks above a certain threshold, roughly that of "I don't need to worry about food or bills".


It's worth noting that while the curve flattens above a threshold, it doesn't level off completely at that threshold, there is still a positive correlation, just a smaller one.


No, that study was constantly misreported on. There's a nice correlation all the way up.


And that threshold would set someone in among richest 1 percent in the world.


And when is that exactly? It definitely isn't making (unadjusted for inflation) the $70k that study suggests.

People are happy when they are secure and unhappy when they are insecure. Who can you name is secure in all of their physical, social, mental, spiritual, etc needs right now?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: