Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

>that also makes good use of that full window?!

To be honest the "Needle In A Haystack" test is the most trivial test for a model that relies on full attention, it's expected to be easy to pass if the model was trained correctly.



i just hope people don't claim "X model support Y context window", when the evaluation is done on "Needle in a haystack" only. It creates so much unnecessary hype.


I agree. I personally don't have high hopes for the 0.5B model.

Phi-2 was 2.7B and it was already regularly outputting complete nonsense.

I ran the 0.5B model of the previous Qwen version (1.5) and it reminded me of one of those lorum ipsum word generators.

The other new Qwen models (7B and up) look good though.


Phi-2 wasn't instruct/chat finetuned and it was very upfront about this, "I tried Phi-2 and it was bad" is a dilletante filter




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: