EleutherAI/lm-evaluation-harness

Add long context evaluation benchmarks such as LongBench and LEval.

Open

#2,180 opened on Aug 5, 2024

View on GitHub
 (3 comments) (1 reaction) (0 assignees)Python (3,306 forks)auto 404
feature requesthelp wanted

Repository metrics

Stars
 (12,755 stars)
PR merge metrics
 (Avg merge 15d 7h) (11 merged PRs in 30d)

Description

Add long context evaluation benchmarks such as LongBench and LEval.

Contributor guide