VEP-bench
VEP-bench is a public benchmark of language models' native ability to predict genetic variant effects. Models answer without internet access or tools, and every response and deterministic score can be inspected.
Leaderboard
Score by cost and token usage
Each line connects evaluated configurations from the same model family. Use the selector to compare the selected task's score with total run cost or total token usage.
Unscored model attempts
An attempt is reported here benchmark-wide when a refusal or content filter in any task prevents a complete, rankable model result. These attempts remain visible regardless of the task selected above and are not included in the leaderboard.