Hacker Newsnew | past | comments | ask | show | jobs | submit | balefulboy's commentslogin

it seems the older models were capped at 10kusd for the runs though?


Damn, that beam of light was a flashbang. I wouldn't call this tasteful UI design, but maybe I just need to go to sleep.


Greentext is eh. Very formulaic, in fact very similar to the bottomless pit one, which I'd argue is better because of it's absurdity. I have to ask, did you mention the older GPT version to fable in the prompt?


Of course I did! Wouldn't be faithfully mediocre without the right context


i got a 502


METR's time horizon is not a reliable metric of LLM capability growth: https://www.transformernews.ai/p/against-the-metr-graph-codi...


Yes I've seen this before, and while the critiques are fair and high quality (and unfortunately not unique to METR) we're missing the forest for the trees here.

First of all, if you take the articles critiques and work out the implications on the METR graph, all you're doing is shifting the curve up or down, it doesn't change the fact that progress is scaling exponentially. While it is technically possible the universe could be throwing a massive pathological curveball to change the conclusion from METR data (which is we've been seeing exponential growth over the last 6 years), I think that seems very far from likely. The fact that we see the same behavior from a variety of sources over a wide variety of tasks and domains is a pretty clear indication that METR while certainly far from perfect is actually painting a consistent picture at least in terms of the rate of progress.

You can look at ECI for a summary benchmark statistic, which does NOT use METR's benchmark, and you see a similar trend. Same with SWE-bench where the task distribution is far more in domain for real world problems. It is a bummer that this METR data can't be better funded. It would probably take $1M or so to really beef it up properly which any of these labs probably have in their couch cushions.


Wow. This deserves to be much more widely read. Thank you for this.


yeah man this sucks. i genuinely do not know how people find this stuff appealing


AI psychosis


I still don't know why people are saying this. I don't really code but from what I've heard on here the models haven't improved since Opus 4.5


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: