enhancementhelp wanted
Repository metrics
- Stars
- (16,939 stars)
- PR merge metrics
- (PR metrics pending)
Description
Currently, users of deepeval can only create their own evaluation dataset/test cases. To support more users fine-tuning their model, deepeval should be able to import standard benchmarks such as MMLU, hellaswag, TruthfulQA, Big Bench, etc.
Comment or DM on discord to discuss more: https://discord.com/invite/a3K9c8GRGt