Reconstruction from compressed evidence
Audio, timbre, harmony, lyrics, dynamics and performance cues are removed. The model receives only a short textual trace of the melody.
MeloText-ID · Symbolic melody identification benchmark
MeloText-ID evaluates whether a language model can identify a familiar melody from a short monophonic sequence represented only as note names and, in one condition, duration values. The task requires the model to recover musically relevant relations from compressed evidence and preserve melodic identity when rhythm is withheld or absolute pitch is changed, while selecting among 25 candidate titles without audio, tools, metadata or cross-request context.
This is a behavioral test of cross-representational reconstruction. It does not identify the internal mechanism or imply subjective auditory experience.
Which melody is encoded?
C C G G A A G | F F E E D D C
Audio, timbre, harmony, lyrics, dynamics and performance cues are removed. The model receives only a short textual trace of the melody.
Transposition preserves interval relations and contour while changing every absolute pitch; the notes-only condition additionally removes duration.
Candidate positions are balanced, choice IDs are graded exactly, and every model completes ten independently randomized repetitions at a fixed 4% chance level.
Human reference: the task is hypothesized to be near-ceiling for the prespecified reference class of skilled musicians who understand the encoding and have previously performed the tested repertoire. No matched human cohort has yet been measured.
Interactive results
The official score includes all four conditions. Any non-empty subset can be inspected as a diagnostic view of sensitivity to transposition and rhythmic information.
| Rank | Model | Score | vs chance | 10-run range | Cost | Details |
|---|---|---|---|---|---|---|
| #1 | Gemini 3.6 Flashgoogle/gemini-3.6-flash | 41.36% | 10.3× | 37–43% | $2.75 | 10 runsProvider pin: google-ai-studio Arithmetic mean: 41.40%
0 parse failures · 0 infrastructure failures |
| #2 | GPT-5.6 Solopenai/gpt-5.6-sol | 40.50% | 10.1× | 36–45% | $13.15 | 10 runsProvider pin: openai Arithmetic mean: 40.60%
0 parse failures · 0 infrastructure failures |
| #3 | Claude Opus 4.8anthropic/claude-opus-4.8 | 35.36% | 8.8× | 34–40% | $5.42 | 10 runsProvider pin: anthropic Arithmetic mean: 35.40%
0 parse failures · 0 infrastructure failures |
| #4 | Kimi K3moonshotai/kimi-k3 | 31.59% | 7.9× | 29–37% | $9.33 | 10 runsProvider pin: moonshotai/int4 Arithmetic mean: 31.70%
34 parse failures · 0 infrastructure failures |
| #5 | Claude Sonnet 5anthropic/claude-sonnet-5 | 22.51% | 5.6× | 19–26% | $1.90 | 10 runsProvider pin: anthropic Arithmetic mean: 22.60%
0 parse failures · 0 infrastructure failures |
| #6 | GPT-5.6 Lunaopenai/gpt-5.6-luna | 22.34% | 5.6× | 20–26% | $3.85 | 10 runsProvider pin: openai Arithmetic mean: 22.40%
16 parse failures · 0 infrastructure failures |
| #7 | Qwen 3.7 Maxqwen/qwen3.7-max | 20.12% | 5.0× | 18–24% | $1.02 | 10 runsProvider pin: alibaba Arithmetic mean: 20.20%
2 parse failures · 0 infrastructure failures |
| #8 | GLM 5.2z-ai/glm-5.2 | 14.98% | 3.7× | 12–18% | $0.83 | 10 runsProvider pin: z-ai/fp8 Arithmetic mean: 15.10%
0 parse failures · 0 infrastructure failures |
| #9 | DeepSeek V4 Prodeepseek/deepseek-v4-pro | 10.06% | 2.5× | 6–14% | $0.68 | 10 runsProvider pin: ionstream/fp4 Arithmetic mean: 10.40%
0 parse failures · 0 infrastructure failures |
| #10 | MiniMax M3minimax/minimax-m3 | 4.99% | 1.2× | 2–11% | $0.19 | 10 runsProvider pin: minimax/fp8 Arithmetic mean: 5.50%
10 parse failures · 0 infrastructure failures |
Bars use a fixed 0–100% scale. The vertical mark is 4% chance. A zero geometric score can mean one repetition scored zero; no epsilon is added.
Four ways to recognize one tune
Absolute geometric scores. Select a condition cell to move the leaderboard to that slice.
| Model | Overall | Original · Notes only | Original · Pitch + duration | Transposed · Notes only | Transposed · Pitch + duration |
|---|---|---|---|---|---|
| Gemini 3.6 Flash | 41.36% | ||||
| GPT-5.6 Sol | 40.50% | ||||
| Claude Opus 4.8 | 35.36% | ||||
| Kimi K3 | 31.59% | ||||
| Claude Sonnet 5 | 22.51% | ||||
| GPT-5.6 Luna | 22.34% | ||||
| Qwen 3.7 Max | 20.12% | ||||
| GLM 5.2 | 14.98% | ||||
| DeepSeek V4 Pro | 10.06% | ||||
| MiniMax M3 | 4.99% |
Darker cells indicate higher scores · every cell is labeled · chance is 4%
Why it matters
Human recognition of familiar melodies is strongly relational: a tune can remain identifiable across different instruments, keys, tempi and notational forms even though its immediate sensory surface changes.
For a language model, the surface input in MeloText-ID is a sequence of tokens. Successful behavior is consistent with organizing those tokens into a relative pitch–time trajectory and connecting that structure to long-term semantic knowledge of a named piece. Exact source-string matching is less sufficient after transposition changes every pitch token, while the notes-only condition removes duration information that would otherwise disambiguate rhythmic structure.
This makes the benchmark a narrow probe of transformation invariance and sparse-cue retrieval—behaviors that are relevant to systematic generalization but do not isolate it. The model sees a single short, lossy description, receives no worked examples or cross-request context, and must reuse the same candidate identity across four controlled representations. Performance that persists across those transformations is evidence that the stable relational structure remains usable despite changes to key, rhythmic detail and textual surface form.
Research on human concept learning shows why this distinction matters. People can often generalize from very few observations when they infer a structured, reusable representation rather than memorize each surface form independently. MeloText-ID applies that concern to musical retrieval: can knowledge acquired during pretraining be made usable from one compressed cue and carried across a transformation? Here, data efficiency refers to the evidence supplied at inference time, not to how many musical examples the model encountered during training.
The relevant object is an ordered configuration of intervals, contour and, when available, duration—not an isolated list of pitch labels.
Each request supplies one excerpt and no demonstrations. The model must recover a candidate identity from substantially less information than an audio recording provides.
The same musical identity must remain accessible in the original and transposed keys, with and without explicit rhythmic values.
These studies motivate the behavioral construct and the relationship between structured representation and generalization. They do not establish how any evaluated model performs the task.
From score to controlled test
The pilot converts a fixed repertoire into four matched stimulus conditions, balances the closed-set answer space, and aggregates repeated measurements without a model-based judge.
Choose 25 highly familiar, nameable melodies with clear monophonic motifs.
Extract the most recognizable passage and check the transcription against score sources.
Create notes-only and duration-prefixed strings, using 32nd-note units for timing.
Transpose the key while preserving intervals, note attacks, durations and measure boundaries.
Present all 25 titles and make every answer position correct exactly once per condition and repetition.
Run ten independently randomized 100-item repetitions, grade exact choice IDs and take their geometric mean.
Notes onlyC C G G A A G
Pitch + duration8C 8C 8G 8G 8A 8A 16G
Every stateless request contains one symbolic melody and all 25 candidate titles. Models receive no audio, browsing, tools, metadata, or cross-request context. `R` means rest; accidentals and octave markers are encoded explicitly.
Download the frozen prompt ↓Candidate permutations are deterministic. Within each condition and repetition, every answer position is correct exactly once. Exact JSON choice IDs are graded automatically; malformed fully received answers remain incorrect.
Within each repetition, accuracy is arithmetic. Across ten repetitions, the score is the exact geometric mean. For a custom slice, selected cells are pooled inside each repetition before taking the geometric mean. No epsilon is added.
score = (A₁ × A₂ × … × A₁₀)^(1/10)
Exact model/provider routes were pinned and fallback was disabled. The complete scored panel cost $39.111479. Hash, route, balance, scoring, export and cost checks passed.
Version 0.1.3-pilot · finalized July 23, 2026 · source run live-20260722T122046Z-a0d2fe0276.
Experiment fingerprint: a0d2fe02767f66b48a6cce5a993698b44786efb78aab917b27f73225a2ef385c
The recognition set
The pilot includes classical themes, national and traditional songs, and widely circulated popular melodies selected for nameability and reproducible monophonic encoding. It is a deliberately compact recognition set, not a balanced survey of global musical culture.
Read the result precisely
MeloText-ID measures closed-set top-1 identification of familiar melodies under controlled changes to pitch and duration encoding. Performance indicates whether melody identity remains accessible from sparse symbolic evidence; it does not establish how the model performs that mapping.
The small set is weighted toward Western classical and globally circulated popular music. Cultural familiarity and musical training matter, and note sequences or close variants may have appeared in pretraining data. Transposition weakens exact surface matching but cannot eliminate memorized transcriptions or training contamination.
Monophonic text removes harmony, timbre, dynamics, articulation and performance. Closed-set recognition is not melody generation or general musical ability.
The public ranking contains only complete, audited evaluations. Additional models will be benchmarked as model access and evaluation capacity expand, with future panels separately versioned so the present result remains reproducible.
The all-condition leaderboard is the official result. Every subset ranking is a transparent diagnostic view and should not replace the preregistered score. The ten repetitions measure sensitivity to balanced answer order and run-to-run variation; they are not population-level confidence intervals.
For the intended human reference class—a skilled musician fluent in this encoding and tested on melodies they have previously performed—the expected ceiling is close to 100% on this 25-way task. Dowling and Fujitani (1971) ↗ reported 99% identification of intact familiar tunes in a five-choice auditory task, although that paradigm supplied audio and was not matched to MeloText-ID. The human ceiling here therefore remains a prospective competence expectation, not a measured baseline. A matched study should control notation fluency, item-level repertoire familiarity and auditory-imagery ability.
The benchmark does not show that a model learned music from few pretraining examples. It measures evidence efficiency at inference time: each decision must be made from one compressed excerpt, without demonstrations, tools, audio or cross-request memory.
Open aggregate evidence
The downloadable record includes aggregate scores, repeated-measure results, the evaluation protocol and integrity commitments. Raw answers and the answer key remain withheld to preserve future evaluation validity.
Independent research
MeloText-ID was conceived and created as an independent study of how AI systems reconstruct relational structure from compressed symbolic evidence and connect it to a stable semantic identity.
Visit alejandrozarzuelo.com ↗