View source

VEP-bench

VEP-bench is a public benchmark of language models' native ability to predict genetic variant effects. Models answer without internet access or tools, and every response and deterministic score can be inspected.

Leaderboard

Score by cost and token usage

Each line connects evaluated configurations from the same model family. Use the selector to compare the selected task's score with total run cost or total token usage.

Unscored model attempts

An attempt is reported here benchmark-wide when a refusal or content filter in any task prevents a complete, rankable model result. These attempts remain visible regardless of the task selected above and are not included in the leaderboard.