Hacker Newsnew | past | comments | ask | show | jobs | submit | megavon's commentslogin

Having played with it for like 3 hours now....I'm probably moving from CC to this


I am about 1 hour into using it with pi.dev. Do you have thinking on high? It is doing good but at one point i had to stop it and say 'you're overthinking this' haha


Yes full send mode on thinking. I have moved on from watching my agents and I don't really care how it thinks. I look at the end result and so far this thing has been blowing me away. No way this is as good as it is this small and fast. Outside Fable, this might be the best thing I've ever used.


What quant are you using?



One hour in, no more Codex for me. This thing rips.


This is INSANE. How did they do this?


"What we've done in this model is not necessarily add more intelligence, but improve the behaviors that lead to a more capable model: more verification, less taking things for granted, not declaring victory early, and being more persistent.”

+

https://poolside.ai/assets/laguna/laguna-m1-xs2-technical-re...


Almost like a built-in heavyweight harness.


Are all AI labs in Google's weight class in crawling and ranking the Web's content? I know OpenAI has contractor subject matter experts in all topics.


"It went from the start of training to launch in under nine weeks..."

This is pretty impressive.


Need to look at SWEBench-Pro, it's super competitive. Suspect they'll catch up given the longer-tail on TB scores.


Just by the (lack of) inter-model variance, I don't think SWEBench-Pro does a very good job of representing model capability. Terminal-Bench seems more challenging and separates the wheat from the chaff.

Also, *ops work, which in my experience can actually be more complicated than SWE is underrepresented there obviously.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: