Read the ZeroScrolls paper to understand the benchmark tasks, then explore the lm evaluation harness repository to understand how existing benchmarks are implemented, and follow the pattern to add a new task.
Tech stack
python
Domain
machine learningai
Issue type
Feature
DifficultyEstimated implementation difficulty for a new contributor, from 1 for very small changes to 5 for expert-level work.
3
Estimated timeA rough time range for an experienced contributor to investigate, implement, test, and prepare a pull request.
1-2 days
Activity statusHow available the issue appears right now: fresh, active, stale, blocked, or waiting on maintainer input.
Fresh
ClarityHow clearly the issue explains the expected change, acceptance criteria, and next step.
Unclear
Prerequisites
PythonGit
Newbie friendlinessA 1-100 score estimating how approachable this issue is for first-time contributors.
40
Daily Newsletter
Get fresh easy issues in your inbox.
Subscribe to GoodFirstIssue Daily for newly found easy issues that are ready for beginner-friendly open source work.