EleutherAI/lm-evaluation-harness

Plans to add new long-context benchmark like LongBench v2, Babilong, InfiniteBench datasets?

Open

#3,224 opened on Aug 8, 2025

View on GitHub
 (7 comments) (1 reaction) (0 assignees)Python (3,306 forks)auto 404
good first issuehelp wanted

Repository metrics

Stars
 (12,755 stars)
PR merge metrics
 (Avg merge 15d 7h) (11 merged PRs in 30d)

Description

Hi lm-eval team, I am wondering if there are plans to add LongBench v2, Babilong, InfiniteBench, and Phonebook datasets to the evaluation tasks? These are useful for long-context LLM evaluation.

Thanks!

Contributor guide