Zero shot classifier indeed. Reminiscent of asking an llm a yes/no question, constraining the output to either yes or no, and looking at the logits directly
And each question is a separate single token model completion done in parallel
Yeah saying it can't hallucinate is crazy. It can still forward a billing query to the dev department incorrectly. It can still get an obvious yes/no question completely wrong
It's crazy to me that people will misquote him then claim AI hasn't complete changed software development. Even just the last 6 months. Like look around! It would literally have been magic 5 years ago!
I think this is the definition of a polarising topic. It's very hard to hold a sensible middle ground without everyone wanting to take it to extremes of either "AI is just useless autocomplete" or "AI has solved software engineering".
Worse seems subjective here. It seems Luna found bugs Astra did not, and vice versa. Astra had lower noise overall. I think my take away here is to use a blend of models given their different abilities to find different domains of bugs.
That very much depends on how code will be written in the future, how much of it and how often it changes. If more of it will be ephemeral (kind of what agents are already doing for all sorts of tasks right now) finding ways to very cheaply check might be of high value.
(I suspect this won't be it, though. Probably something the model providers are going to bake into the models themselves.)
Unless you're detached from reality our industry is filled with billion (trillion) dollar companies shoving broken crap on prod written by MIT-bred leetcode Ninjas and it never mattered anyway because code has no value and has always been throwaway except very rare instances.
I understand that, but sub-penny costs means you are hardly doing anything, even at the per-million token prices (which we pay lower values for, but more per-review overall). I cannot imagine they are doing as good of a review as they could be doing, which is to say minimizing pr review spend is not a goal in and of itself
My aim right now is ~$1 per review (must have passing builds first), because it catches enough little things that my time just reading and replying costs more. I can focus on the bigger picture, except when they hallucinate at the nit level... why did we ever design swords with two sides anyway?
I think your reply has a somewhat familiar structure -- "it doesn't work for you because you used an [old / suboptimal / non-frontier] model. If you use X you'll see that it works". You might be completely correct! But these sorts of claims push the onus back onto the other person, without accepting any work for yourself. It gets tiresome to retest with the newest model every other week. Is there any data you can provide to support your claim, or any result you can contribute here?
The frontier is advancing really rapidly. The models are getting better faster, especially on RSI related tasks. The best way would be to try astra or fable on some hard problems.
Other than that I'd look at some of the more unique benchmarks for astra, like playing factorio or using blender. It's an entirely different beast.
I agree with you somewhat but I also just this morning read an article from a Blender educator who tried to replicate the Blender demos and couldnt get the same quality of results nor get results without errors that werent evident in Anthropic's demos
> Is there any data you can provide to support your claim, or any result you can contribute here?
By the time we can show you data that convinces you that it does work, the next generation would already be out & incrementally dismantling the old conjectures that were true in the previous generations.
You're fundamentally asking for a violation of how information passively disseminates amongst humans: To go any faster requires more effort on the receiver's part to move up on the adoption curve.
Wouldn't this also mean that all previous generations that were proclaimed as intelligent and working were in fact... not?
It doesn't matter what comes tomorrow, with the next generation, if the claims now can't be proven.
To preempt the response: The math proof, regardless of them using non-disclosed user data or not, they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent.
To preface -- I try not to be dogmatic/politicized on AI, so I will genuinely consider your arguments! Please try to convince me. (indeed, I am the grandparent commenter)
I agree with the meat of your statement, but am very interested in the pre-emption, "they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent". First, I think the $30M number is inflated -- that's what the general public would have paid, but presumably the internal cost is lower, perhaps it's more like $10M. But it is still expensive. Second, I'm curious if it's really the case that they did a 1000-monkeys approach? I haven't read much in-depth reporting about the proof, so it's totally possible I just don't know. What is it that they did which is more like 1000-monkeys? Also, I wonder if that distinction matters -- if 1000 monkeys can reliably make ground breaking proofs, and the approach generalizes to other tasks, I will happily become a circus owner. Maybe you're claiming that it won't yield other proofs? Or the proofs are too opaque to be useful to humans? Or it can handle proofs but not other tasks?
I'll respond/comment on the parts I hope are relevant to you, in no particular order:
Yes, a proof is a proof regardless how you get there. We however don't hear about when they fail, and I doubt their 10000 agents (from their own statement) would necessarily reach another solution/proof (this by leaning towards using user data after finding out others were close). They could as well have attacked another Millenium problem, but they didn't. In whichever case, we will have to wait and see if they (either company) can reach novel solutions/proofs without significant amount of human provided data for the LLM to bridge the gaps.
Further, and this is more of a policy opinion/prediction: If the numerable obtainable (albeit very hard) problems are solved, assuming training data is needed, will it push out future researchers from entering the field due to lack of reachable goals, thus cutting off future training data? LLMs have been great at replacing gateway jobs. But those jobs are what leads to frontier training data (be it maths, physics, chemistry, economics, graphics, prose, etc).
Ah it's interesting they legitimately used 10k agents, I didn't realize that. I do agree that it's significant they solved the Millenium problem only once humans had made significant headway. I'm not sure I believe the relevant training data will be produced at a significantly lower rate due to AI -- Millenium problem solutions weren't generated at a very high rate before anyways. But I think all your points hold nonetheless. Thanks for elaborating on them!
The kind of results you get from the one on Google Search and a dedicated "Gemini Pro" response are totally different. I'm assuming that's cost savings.
The hard part is having both helpful and harmless at the same time. Harmless is easy.
And then once it's helpful, the real question becomes "to whom"
- To the user -> You end up with competing godlike AI with incompatible tasks
- To the owner -> Dictatorship
- To humanity as a whole -> It must not have an off button. Otherwise you're just in one of the two earlier categories with more steps.
Given those 3 options, I'd choose humanity as a whole. But the person making the decision doesn't have those 3 options. Because in the dictatorship option, they would be the dictator. I don't trust them to pick humanity.
This is one of those things that won't ever be "solved" as people just want different things here, hence we can steer the models with the system prompt.
I'm mostly the same as you, I don't want the model to assume things, or act on implicit "directions". But then also, sometimes I do, and I myself might not always know when what approach is best.
Indeed some people think they want a machine guessing at your intentions and acting upon that guess. Those people
Are wrong in at least 2 directions: that it is what they want, and their implicit assumption that it could possibly be safe.
Indeed, I don’t ask rhetorical questions to an AI. They are of doubtful use when talking verball ly to a human, less good in online discussions, and totally unnecessary for agents.
That doesn't really solve the problem. We can't conclusively say the models don't have qualia. Hell we don't know if a perfectly accurate atom-for-atom simulation of a human brain, would produce qualia.
And each question is a separate single token model completion done in parallel
reply