Skip to content

LLM Benchmark Hub

Model evidence, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
LLM Benchmark Hub
Loading data...

VSI-Super-Wild — leaderboard

Metric: Overall Accuracy (%). Source: vsi-super-wild.github.io. 14 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)44.36
2GPT-5.434.52
3Qwen 3.5 9B33.87
4InternVL3.5-8B32.18
5Qwen 2 VL 7B26.67
6Gemini 3.1 Flash Lite23.85

Interactive version: aibenchmarks.dev/benchmark?slug=vsi-super-wild · How the rankings work · Data refreshed daily, snapshot 2026-07-20.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.