Hacker Newsnew | past | comments | ask | show | jobs | submit | MisterBiggs's commentslogin

Yup! I feel pretty strongly that every little nit pick and instruction you pass into your model is murdering your output. Having a hook that executes on tool calls is significantly better than telling your agent to follow your repos specific format/lint/style/test constraints


Skills for repitition are totally valid. Having a version control skill that explains that I use gitea works great. My point is that asking for a skill that tells us if our program will get stuck before taking on a halting problem won't get you any further than just starting the task with xhigh thinking


This is a pretty neat, I suspect that eventually every skill will have some sort of validation/verification loop like this


What happens once an agent can reliably get 100% on swebench?


For writing with an LLM you really have to be precise and honest with your AI usage and I don't think the disclosure here does a good job at that. If you sample a few random pages there are huge shifts in style and tone.

The future where your AI expands your sentence into a few paragraphs that my AI distills down into a sentence sucks just send me your rough draft


I'd rather people bounce back and forward with the AI until they're happy, but then don't fluff it up and expand it needlessly in the final step, spit out a condensed, tight bullet list, no fluff no purple prose no emojis and send me that instead.


It has nothing to do with running models locally, its perfect because its incredibly cheap, capable, small, and quiet.


I was really hoping for a new form factor or new killer features. Its too bad that the general public can't behave themselves with simple tech like this


I feel like I made a mistake going with subdomains for each project instead of root folders, but here is my landing page https://ansonbiggs.com/


Cool to see that Anthropic and OpenAI are really diverging in their target customer. Anthropic seems much more focused on making tools for pros while OpenAI is aiming at consumers.


Shield AI | Autonomy Applications Manager | Full Time | ONSITE Washington DC

Shield AI's mission is to protect service members and civilians by creating intelligent systems.

My team needs a manager, its a very well positioned team at a rapidly growing company working with some really awesome people. Any autonomy or robotics experience is a huge plus https://jobs.lever.co/shieldai/7952271c-8578-42d8-a587-4707c...

Contact info on my profile if you have any questions I'm happy to chat!


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: