confident-ai/deepeval

Support popular benchmarks

Open

#508 opened on Feb 22, 2024

View on GitHub
 (9 comments) (0 reactions) (0 assignees)Python (1,677 forks)auto 404
enhancementhelp wanted

Repository metrics

Stars
 (16,939 stars)
PR merge metrics
 (PR metrics pending)

Description

Currently, users of deepeval can only create their own evaluation dataset/test cases. To support more users fine-tuning their model, deepeval should be able to import standard benchmarks such as MMLU, hellaswag, TruthfulQA, Big Bench, etc.

Comment or DM on discord to discuss more: https://discord.com/invite/a3K9c8GRGt

Contributor guide