To be honest the "Needle In A Haystack" test is the most trivial test for a model that relies on full attention, it's expected to be easy to pass if the model was trained correctly.
i just hope people don't claim "X model support Y context window", when the evaluation is done on "Needle in a haystack" only. It creates so much unnecessary hype.
To be honest the "Needle In A Haystack" test is the most trivial test for a model that relies on full attention, it's expected to be easy to pass if the model was trained correctly.