MeloText-ID · Symbolic melody identification benchmark

Can AI hear a melody from text?

MeloText-ID evaluates whether a language model can identify a familiar melody from a short monophonic sequence represented only as note names and, in one condition, duration values. The task requires the model to recover musically relevant relations from compressed evidence and preserve melodic identity when rhythm is withheld or absolute pitch is changed, while selecting among 25 candidate titles without audio, tools, metadata or cross-request context.

This is a behavioral test of cross-representational reconstruction. It does not identify the internal mechanism or imply subjective auditory experience.

Held-out demonstration · excluded from evaluation

Which melody is encoded?

C C G G A A G  |  F F E E D D C

01

Reconstruction from compressed evidence

Audio, timbre, harmony, lyrics, dynamics and performance cues are removed. The model receives only a short textual trace of the melody.

02

Invariance under controlled transformation

Transposition preserves interval relations and contour while changing every absolute pitch; the notes-only condition additionally removes duration.

03

Reproducible closed-set measurement

Candidate positions are balanced, choice IDs are graded exactly, and every model completes ten independently randomized repetitions at a fixed 4% chance level.

Highest official score41.36%Gemini 3.6 Flash
Closed-set chance4%1 of 25 candidates
Best / chance10.3×chance-level accuracy
10,000 scored decisions10 complete models25 melodies × 4 conditions Audit passed

Human reference: the task is hypothesized to be near-ceiling for the prespecified reference class of skilled musicians who understand the encoding and have previously performed the tested repertoire. No matched human cohort has yet been measured.

Interactive results

Compare performance across controlled conditions.

The official score includes all four conditions. Any non-empty subset can be inspected as a diagnostic view of sensitivity to transposition and rhythmic information.

Official

All four conditions

Geometric mean across ten complete repetitions · chance = 4%

MeloText-ID model leaderboard for all four conditions
RankModelScorevs chance10-run rangeCostDetails
#1Gemini 3.6 Flashgoogle/gemini-3.6-flash41.36%10.3×3743%$2.75
10 runs

Provider pin: google-ai-studio

Arithmetic mean: 41.40%

  1. R137.00%
  2. R242.00%
  3. R342.00%
  4. R442.00%
  5. R542.00%
  6. R643.00%
  7. R743.00%
  8. R839.00%
  9. R943.00%
  10. R1041.00%

0 parse failures · 0 infrastructure failures

#2GPT-5.6 Solopenai/gpt-5.6-sol40.50%10.1×3645%$13.15
10 runs

Provider pin: openai

Arithmetic mean: 40.60%

  1. R137.00%
  2. R244.00%
  3. R338.00%
  4. R442.00%
  5. R545.00%
  6. R640.00%
  7. R740.00%
  8. R843.00%
  9. R936.00%
  10. R1041.00%

0 parse failures · 0 infrastructure failures

#3Claude Opus 4.8anthropic/claude-opus-4.835.36%8.8×3440%$5.42
10 runs

Provider pin: anthropic

Arithmetic mean: 35.40%

  1. R140.00%
  2. R235.00%
  3. R335.00%
  4. R434.00%
  5. R534.00%
  6. R634.00%
  7. R736.00%
  8. R837.00%
  9. R935.00%
  10. R1034.00%

0 parse failures · 0 infrastructure failures

#4Kimi K3moonshotai/kimi-k331.59%7.9×2937%$9.33
10 runs

Provider pin: moonshotai/int4

Arithmetic mean: 31.70%

  1. R131.00%
  2. R229.00%
  3. R336.00%
  4. R429.00%
  5. R531.00%
  6. R629.00%
  7. R730.00%
  8. R833.00%
  9. R937.00%
  10. R1032.00%

34 parse failures · 0 infrastructure failures

#5Claude Sonnet 5anthropic/claude-sonnet-522.51%5.6×1926%$1.90
10 runs

Provider pin: anthropic

Arithmetic mean: 22.60%

  1. R119.00%
  2. R222.00%
  3. R322.00%
  4. R421.00%
  5. R526.00%
  6. R622.00%
  7. R723.00%
  8. R825.00%
  9. R921.00%
  10. R1025.00%

0 parse failures · 0 infrastructure failures

#6GPT-5.6 Lunaopenai/gpt-5.6-luna22.34%5.6×2026%$3.85
10 runs

Provider pin: openai

Arithmetic mean: 22.40%

  1. R121.00%
  2. R224.00%
  3. R326.00%
  4. R422.00%
  5. R520.00%
  6. R621.00%
  7. R723.00%
  8. R821.00%
  9. R923.00%
  10. R1023.00%

16 parse failures · 0 infrastructure failures

#7Qwen 3.7 Maxqwen/qwen3.7-max20.12%5.0×1824%$1.02
10 runs

Provider pin: alibaba

Arithmetic mean: 20.20%

  1. R121.00%
  2. R219.00%
  3. R319.00%
  4. R419.00%
  5. R523.00%
  6. R620.00%
  7. R720.00%
  8. R819.00%
  9. R918.00%
  10. R1024.00%

2 parse failures · 0 infrastructure failures

#8GLM 5.2z-ai/glm-5.214.98%3.7×1218%$0.83
10 runs

Provider pin: z-ai/fp8

Arithmetic mean: 15.10%

  1. R112.00%
  2. R215.00%
  3. R312.00%
  4. R416.00%
  5. R515.00%
  6. R616.00%
  7. R717.00%
  8. R818.00%
  9. R914.00%
  10. R1016.00%

0 parse failures · 0 infrastructure failures

#9DeepSeek V4 Prodeepseek/deepseek-v4-pro10.06%2.5×614%$0.68
10 runs

Provider pin: ionstream/fp4

Arithmetic mean: 10.40%

  1. R112.00%
  2. R26.00%
  3. R39.00%
  4. R49.00%
  5. R514.00%
  6. R612.00%
  7. R78.00%
  8. R813.00%
  9. R98.00%
  10. R1013.00%

0 parse failures · 0 infrastructure failures

#10MiniMax M3minimax/minimax-m34.99%1.2×211%$0.19
10 runs

Provider pin: minimax/fp8

Arithmetic mean: 5.50%

  1. R14.00%
  2. R28.00%
  3. R35.00%
  4. R43.00%
  5. R56.00%
  6. R65.00%
  7. R76.00%
  8. R82.00%
  9. R911.00%
  10. R105.00%

10 parse failures · 0 infrastructure failures

Bars use a fixed 0–100% scale. The vertical mark is 4% chance. A zero geometric score can mean one repetition scored zero; no epsilon is added.

Four ways to recognize one tune

Condition matrix

Absolute geometric scores. Select a condition cell to move the leaderboard to that slice.

ModelOverallOriginal · Notes onlyOriginal · Pitch + durationTransposed · Notes onlyTransposed · Pitch + duration
Gemini 3.6 Flash41.36%
GPT-5.6 Sol40.50%
Claude Opus 4.835.36%
Kimi K331.59%
Claude Sonnet 522.51%
GPT-5.6 Luna22.34%
Qwen 3.7 Max20.12%
GLM 5.214.98%
DeepSeek V4 Pro10.06%
MiniMax M34.99%

Darker cells indicate higher scores · every cell is labeled · chance is 4%

Why it matters

Musical identity is relational, not merely textual.

Human recognition of familiar melodies is strongly relational: a tune can remain identifiable across different instruments, keys, tempi and notational forms even though its immediate sensory surface changes.

For a language model, the surface input in MeloText-ID is a sequence of tokens. Successful behavior is consistent with organizing those tokens into a relative pitch–time trajectory and connecting that structure to long-term semantic knowledge of a named piece. Exact source-string matching is less sufficient after transposition changes every pitch token, while the notes-only condition removes duration information that would otherwise disambiguate rhythmic structure.

This makes the benchmark a narrow probe of transformation invariance and sparse-cue retrieval—behaviors that are relevant to systematic generalization but do not isolate it. The model sees a single short, lossy description, receives no worked examples or cross-request context, and must reuse the same candidate identity across four controlled representations. Performance that persists across those transformations is evidence that the stable relational structure remains usable despite changes to key, rhythmic detail and textual surface form.

Research on human concept learning shows why this distinction matters. People can often generalize from very few observations when they infer a structured, reusable representation rather than memorize each surface form independently. MeloText-ID applies that concern to musical retrieval: can knowledge acquired during pretraining be made usable from one compressed cue and carried across a transformation? Here, data efficiency refers to the evidence supplied at inference time, not to how many musical examples the model encountered during training.

01

Relational abstraction

The relevant object is an ordered configuration of intervals, contour and, when available, duration—not an isolated list of pitch labels.

02

Sparse-evidence retrieval

Each request supplies one excerpt and no demonstrations. The model must recover a candidate identity from substantially less information than an audio recording provides.

03

Systematic reuse

The same musical identity must remain accessible in the original and transposed keys, with and without explicit rhythmic values.

Research context and primary literature

These studies motivate the behavioral construct and the relationship between structured representation and generalization. They do not establish how any evaluated model performs the task.

From score to controlled test

How the benchmark was built

The pilot converts a fixed repertoire into four matched stimulus conditions, balances the closed-set answer space, and aggregates repeated measurements without a model-based judge.

  1. 01

    Select repertoire

    Choose 25 highly familiar, nameable melodies with clear monophonic motifs.

  2. 02

    Verify motifs

    Extract the most recognizable passage and check the transcription against score sources.

  3. 03

    Construct encodings

    Create notes-only and duration-prefixed strings, using 32nd-note units for timing.

  4. 04

    Apply transformation

    Transpose the key while preserving intervals, note attacks, durations and measure boundaries.

  5. 05

    Balance choices

    Present all 25 titles and make every answer position correct exactly once per condition and repetition.

  6. 06

    Repeat and aggregate

    Run ten independently randomized 100-item repetitions, grade exact choice IDs and take their geometric mean.

One melody, two textual surfaces

Notes onlyC C G G A A G

Pitch + duration8C 8C 8G 8G 8A 8A 16G

Task, prompt and stimulus construction

Every stateless request contains one symbolic melody and all 25 candidate titles. Models receive no audio, browsing, tools, metadata, or cross-request context. `R` means rest; accidentals and octave markers are encoded explicitly.

Download the frozen prompt ↓
Randomization, balance and grading

Candidate permutations are deterministic. Within each condition and repetition, every answer position is correct exactly once. Exact JSON choice IDs are graded automatically; malformed fully received answers remain incorrect.

Scoring and aggregation

Within each repetition, accuracy is arithmetic. Across ten repetitions, the score is the exact geometric mean. For a custom slice, selected cells are pooled inside each repetition before taking the geometric mean. No epsilon is added.

score = (A₁ × A₂ × … × A₁₀)^(1/10)

Model routing, costs and audit

Exact model/provider routes were pinned and fallback was disabled. The complete scored panel cost $39.111479. Hash, route, balance, scoring, export and cost checks passed.

Provenance and reproducibility

Version 0.1.3-pilot · finalized July 23, 2026 · source run live-20260722T122046Z-a0d2fe0276.

Experiment fingerprint: a0d2fe02767f66b48a6cce5a993698b44786efb78aab917b27f73225a2ef385c

The recognition set

A controlled set of 25 familiar melodies.

The pilot includes classical themes, national and traditional songs, and widely circulated popular melodies selected for nameability and reproducible monophonic encoding. It is a deliberately compact recognition set, not a balanced survey of global musical culture.

  1. 011812 Overture
  2. 02Ave Maria (Schubert)
  3. 03Beethoven's Fifth Symphony
  4. 04Ode to Joy (Beethoven's Ninth Symphony)
  5. 05Bella Ciao
  6. 06The Blue Danube
  7. 07Canon in D (Pachelbel)
  8. 08Clair de Lune
  9. 09Eine kleine Nachtmusik
  10. 10Fur Elise
  11. 11Hallelujah Chorus
  12. 12Happy Birthday to You
  13. 13In the Hall of the Mountain King
  14. 14The Internationale
  15. 15La donna e mobile
  16. 16March of the Volunteers (Chinese National Anthem)
  17. 17Mo Li Hua (Jasmine Flower)
  18. 18O Sole Mio
  19. 19Pomp and Circumstance March No. 1
  20. 20Queen of the Night Aria
  21. 21Rondo Alla Turca (Turkish March)
  22. 22The Star-Spangled Banner
  23. 23Toreador Song
  24. 24Two Tigers (Frere Jacques melody)
  25. 25Spring (Vivaldi, The Four Seasons)

Read the result precisely

Provenance & limitations

What this benchmark shows

MeloText-ID measures closed-set top-1 identification of familiar melodies under controlled changes to pitch and duration encoding. Performance indicates whether melody identity remains accessible from sparse symbolic evidence; it does not establish how the model performs that mapping.

Cultural and training-data confounds

The small set is weighted toward Western classical and globally circulated popular music. Cultural familiarity and musical training matter, and note sequences or close variants may have appeared in pretraining data. Transposition weakens exact surface matching but cannot eliminate memorized transcriptions or training contamination.

What the representation omits

Monophonic text removes harmony, timbre, dynamics, articulation and performance. Closed-set recognition is not melody generation or general musical ability.

Panel coverage

The public ranking contains only complete, audited evaluations. Additional models will be benchmarked as model access and evaluation capacity expand, with future panels separately versioned so the present result remains reproducible.

Official versus exploratory

The all-condition leaderboard is the official result. Every subset ranking is a transparent diagnostic view and should not replace the preregistered score. The ten repetitions measure sensitivity to balanced answer order and run-to-run variation; they are not population-level confidence intervals.

Human ceiling and future baseline

For the intended human reference class—a skilled musician fluent in this encoding and tested on melodies they have previously performed—the expected ceiling is close to 100% on this 25-way task. Dowling and Fujitani (1971) ↗ reported 99% identification of intact familiar tunes in a five-choice auditory task, although that paradigm supplied audio and was not matched to MeloText-ID. The human ceiling here therefore remains a prospective competence expectation, not a measured baseline. A matched study should control notation fluency, item-level repertoire familiarity and auditory-imagery ability.

What “few data” means here

The benchmark does not show that a model learned music from few pretraining examples. It measures evidence efficiency at inference time: each decision must be made from one compressed excerpt, without demonstrations, tools, audio or cross-request memory.

Open aggregate evidence

Public evidence and reproducibility artifacts

The downloadable record includes aggregate scores, repeated-measure results, the evaluation protocol and integrity commitments. Raw answers and the answer key remain withheld to preserve future evaluation validity.

Independent research

Created by
Alejandro Zarzuelo.

MeloText-ID was conceived and created as an independent study of how AI systems reconstruct relational structure from compressed symbolic evidence and connect it to a stable semantic identity.

Visit alejandrozarzuelo.com ↗