Hacker Newsnew | past | comments | ask | show | jobs | submit | StevenWaterman's commentslogin

Zero shot classifier indeed. Reminiscent of asking an llm a yes/no question, constraining the output to either yes or no, and looking at the logits directly

And each question is a separate single token model completion done in parallel


Yeah saying it can't hallucinate is crazy. It can still forward a billing query to the dev department incorrectly. It can still get an obvious yes/no question completely wrong

It's crazy to me that people will misquote him then claim AI hasn't complete changed software development. Even just the last 6 months. Like look around! It would literally have been magic 5 years ago!

I think this is the definition of a polarising topic. It's very hard to hold a sensible middle ground without everyone wanting to take it to extremes of either "AI is just useless autocomplete" or "AI has solved software engineering".

$0.10 extra per pr review is nothing. What software company is willing to accept worse reviews and less bugs found to save 10 cents?

Worse seems subjective here. It seems Luna found bugs Astra did not, and vice versa. Astra had lower noise overall. I think my take away here is to use a blend of models given their different abilities to find different domains of bugs.

That very much depends on how code will be written in the future, how much of it and how often it changes. If more of it will be ephemeral (kind of what agents are already doing for all sorts of tasks right now) finding ways to very cheaply check might be of high value.

(I suspect this won't be it, though. Probably something the model providers are going to bake into the models themselves.)


Exactly this. It's still (at this time) cheaper than a developer that would most likely perform worse.

Unless you're detached from reality our industry is filled with billion (trillion) dollar companies shoving broken crap on prod written by MIT-bred leetcode Ninjas and it never mattered anyway because code has no value and has always been throwaway except very rare instances.

Per the article the luna review cost $0.004 and astra cost $0.113. The headline is per million tokens

I understand that, but sub-penny costs means you are hardly doing anything, even at the per-million token prices (which we pay lower values for, but more per-review overall). I cannot imagine they are doing as good of a review as they could be doing, which is to say minimizing pr review spend is not a goal in and of itself

My aim right now is ~$1 per review (must have passing builds first), because it catches enough little things that my time just reading and replying costs more. I can focus on the bigger picture, except when they hallucinate at the nit level... why did we ever design swords with two sides anyway?


As someone who used to use Gemini a lot, if you are predominantly using Gemini you don't know what the current state of things is like

I think your reply has a somewhat familiar structure -- "it doesn't work for you because you used an [old / suboptimal / non-frontier] model. If you use X you'll see that it works". You might be completely correct! But these sorts of claims push the onus back onto the other person, without accepting any work for yourself. It gets tiresome to retest with the newest model every other week. Is there any data you can provide to support your claim, or any result you can contribute here?

The frontier is advancing really rapidly. The models are getting better faster, especially on RSI related tasks. The best way would be to try astra or fable on some hard problems.

Other than that I'd look at some of the more unique benchmarks for astra, like playing factorio or using blender. It's an entirely different beast.


I agree with you somewhat but I also just this morning read an article from a Blender educator who tried to replicate the Blender demos and couldnt get the same quality of results nor get results without errors that werent evident in Anthropic's demos

> Is there any data you can provide to support your claim, or any result you can contribute here?

By the time we can show you data that convinces you that it does work, the next generation would already be out & incrementally dismantling the old conjectures that were true in the previous generations.

You're fundamentally asking for a violation of how information passively disseminates amongst humans: To go any faster requires more effort on the receiver's part to move up on the adoption curve.


Wouldn't this also mean that all previous generations that were proclaimed as intelligent and working were in fact... not?

It doesn't matter what comes tomorrow, with the next generation, if the claims now can't be proven.

To preempt the response: The math proof, regardless of them using non-disclosed user data or not, they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent.


To preface -- I try not to be dogmatic/politicized on AI, so I will genuinely consider your arguments! Please try to convince me. (indeed, I am the grandparent commenter)

I agree with the meat of your statement, but am very interested in the pre-emption, "they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent". First, I think the $30M number is inflated -- that's what the general public would have paid, but presumably the internal cost is lower, perhaps it's more like $10M. But it is still expensive. Second, I'm curious if it's really the case that they did a 1000-monkeys approach? I haven't read much in-depth reporting about the proof, so it's totally possible I just don't know. What is it that they did which is more like 1000-monkeys? Also, I wonder if that distinction matters -- if 1000 monkeys can reliably make ground breaking proofs, and the approach generalizes to other tasks, I will happily become a circus owner. Maybe you're claiming that it won't yield other proofs? Or the proofs are too opaque to be useful to humans? Or it can handle proofs but not other tasks?


I'll respond/comment on the parts I hope are relevant to you, in no particular order:

Yes, a proof is a proof regardless how you get there. We however don't hear about when they fail, and I doubt their 10000 agents (from their own statement) would necessarily reach another solution/proof (this by leaning towards using user data after finding out others were close). They could as well have attacked another Millenium problem, but they didn't. In whichever case, we will have to wait and see if they (either company) can reach novel solutions/proofs without significant amount of human provided data for the LLM to bridge the gaps.

Further, and this is more of a policy opinion/prediction: If the numerable obtainable (albeit very hard) problems are solved, assuming training data is needed, will it push out future researchers from entering the field due to lack of reachable goals, thus cutting off future training data? LLMs have been great at replacing gateway jobs. But those jobs are what leads to frontier training data (be it maths, physics, chemistry, economics, graphics, prose, etc).


Ah it's interesting they legitimately used 10k agents, I didn't realize that. I do agree that it's significant they solved the Millenium problem only once humans had made significant headway. I'm not sure I believe the relevant training data will be produced at a significantly lower rate due to AI -- Millenium problem solutions weren't generated at a very high rate before anyways. But I think all your points hold nonetheless. Thanks for elaborating on them!

Maybe Gemini is loosey goosey on purpose so we angrily correct it - then feed something on the back end that trains a different model?

It's so abysmally bad on Google search... and it's free. Isn't Google the great pioneer of the product is us?


The kind of results you get from the one on Google Search and a dedicated "Gemini Pro" response are totally different. I'm assuming that's cost savings.

> You cannot prevent (2) via any alignment process

A little bit too categorical. GOODY-2 wouldn't do it. https://www.goody2.ai/

The hard part is having both helpful and harmless at the same time. Harmless is easy.

And then once it's helpful, the real question becomes "to whom"

- To the user -> You end up with competing godlike AI with incompatible tasks

- To the owner -> Dictatorship

- To humanity as a whole -> It must not have an off button. Otherwise you're just in one of the two earlier categories with more steps.

Given those 3 options, I'd choose humanity as a whole. But the person making the decision doesn't have those 3 options. Because in the dictatorship option, they would be the dictator. I don't trust them to pick humanity.


i will get back to try gandalfing goody2

> Main problem: the quality dramatically hits the shitter done once it falls back to "pdf2latext" due to complex tables.

Could try screenshotting the PDF and passing that to gemini


This is what I'm doing by now. Uploading the entire .pdf dynamically then giving it a pipe to ask questions.

> For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed.

That's exactly what i want to happen. I hate when it assumes my direct question was an indirect instruction


This is one of those things that won't ever be "solved" as people just want different things here, hence we can steer the models with the system prompt.

I'm mostly the same as you, I don't want the model to assume things, or act on implicit "directions". But then also, sometimes I do, and I myself might not always know when what approach is best.


Indeed some people think they want a machine guessing at your intentions and acting upon that guess. Those people Are wrong in at least 2 directions: that it is what they want, and their implicit assumption that it could possibly be safe.

Indeed, I don’t ask rhetorical questions to an AI. They are of doubtful use when talking verball ly to a human, less good in online discussions, and totally unnecessary for agents.


That doesn't really solve the problem. We can't conclusively say the models don't have qualia. Hell we don't know if a perfectly accurate atom-for-atom simulation of a human brain, would produce qualia.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: