MeloText-ID · Cross-representational AI benchmark

Can an AI hear music in plain text?

A melody enters as symbols. No audio. No tools. Twenty-five possible songs. MeloText-ID measures whether a language model can recover the tune—even when rhythm disappears or every pitch moves to a new key.

“Hear” is shorthand for successful cross-representational inference, not a claim about subjective experience.

Demonstration only · not scored

What do you hear?

C C G G A A G  |  F F E E D D C

01

Symbolic, not sonic

Only textual pitch and duration sequences reach the model.

02

Verified, not judged

Exact choice IDs are graded automatically. No LLM evaluator.

03

Transposed, not copied

Key changes test identity beyond absolute-pitch matching.

Highest official score41.36%Gemini 3.6 Flash
Closed-set chance4%1 of 25 candidates
Best / chance10.3×relative lift
10,000 scored decisions10 complete models25 melodies × 4 conditions Audit passed

Interactive results

Choose what the ranking means.

The official score includes all four conditions. Select any non-empty combination to inspect a different capability slice.

Official

All four conditions

Geometric mean across ten complete repetitions · chance = 4%

MeloText-ID model leaderboard for all four conditions
RankModelScorevs chance10-run rangeCostDetails
#1Gemini 3.6 Flashgoogle/gemini-3.6-flash41.36%10.3×3743%$2.75
10 runs

Provider pin: google-ai-studio

Arithmetic mean: 41.40%

  1. R137.00%
  2. R242.00%
  3. R342.00%
  4. R442.00%
  5. R542.00%
  6. R643.00%
  7. R743.00%
  8. R839.00%
  9. R943.00%
  10. R1041.00%

0 parse failures · 0 infrastructure failures

#2GPT-5.6 Solopenai/gpt-5.6-sol40.50%10.1×3645%$13.15
10 runs

Provider pin: openai

Arithmetic mean: 40.60%

  1. R137.00%
  2. R244.00%
  3. R338.00%
  4. R442.00%
  5. R545.00%
  6. R640.00%
  7. R740.00%
  8. R843.00%
  9. R936.00%
  10. R1041.00%

0 parse failures · 0 infrastructure failures

#3Claude Opus 4.8anthropic/claude-opus-4.835.36%8.8×3440%$5.42
10 runs

Provider pin: anthropic

Arithmetic mean: 35.40%

  1. R140.00%
  2. R235.00%
  3. R335.00%
  4. R434.00%
  5. R534.00%
  6. R634.00%
  7. R736.00%
  8. R837.00%
  9. R935.00%
  10. R1034.00%

0 parse failures · 0 infrastructure failures

#4Kimi K3moonshotai/kimi-k331.59%7.9×2937%$9.33
10 runs

Provider pin: moonshotai/int4

Arithmetic mean: 31.70%

  1. R131.00%
  2. R229.00%
  3. R336.00%
  4. R429.00%
  5. R531.00%
  6. R629.00%
  7. R730.00%
  8. R833.00%
  9. R937.00%
  10. R1032.00%

34 parse failures · 0 infrastructure failures

#5Claude Sonnet 5anthropic/claude-sonnet-522.51%5.6×1926%$1.90
10 runs

Provider pin: anthropic

Arithmetic mean: 22.60%

  1. R119.00%
  2. R222.00%
  3. R322.00%
  4. R421.00%
  5. R526.00%
  6. R622.00%
  7. R723.00%
  8. R825.00%
  9. R921.00%
  10. R1025.00%

0 parse failures · 0 infrastructure failures

#6GPT-5.6 Lunaopenai/gpt-5.6-luna22.34%5.6×2026%$3.85
10 runs

Provider pin: openai

Arithmetic mean: 22.40%

  1. R121.00%
  2. R224.00%
  3. R326.00%
  4. R422.00%
  5. R520.00%
  6. R621.00%
  7. R723.00%
  8. R821.00%
  9. R923.00%
  10. R1023.00%

16 parse failures · 0 infrastructure failures

#7Qwen 3.7 Maxqwen/qwen3.7-max20.12%5.0×1824%$1.02
10 runs

Provider pin: alibaba

Arithmetic mean: 20.20%

  1. R121.00%
  2. R219.00%
  3. R319.00%
  4. R419.00%
  5. R523.00%
  6. R620.00%
  7. R720.00%
  8. R819.00%
  9. R918.00%
  10. R1024.00%

2 parse failures · 0 infrastructure failures

#8GLM 5.2z-ai/glm-5.214.98%3.7×1218%$0.83
10 runs

Provider pin: z-ai/fp8

Arithmetic mean: 15.10%

  1. R112.00%
  2. R215.00%
  3. R312.00%
  4. R416.00%
  5. R515.00%
  6. R616.00%
  7. R717.00%
  8. R818.00%
  9. R914.00%
  10. R1016.00%

0 parse failures · 0 infrastructure failures

#9DeepSeek V4 Prodeepseek/deepseek-v4-pro10.06%2.5×614%$0.68
10 runs

Provider pin: ionstream/fp4

Arithmetic mean: 10.40%

  1. R112.00%
  2. R26.00%
  3. R39.00%
  4. R49.00%
  5. R514.00%
  6. R612.00%
  7. R78.00%
  8. R813.00%
  9. R98.00%
  10. R1013.00%

0 parse failures · 0 infrastructure failures

#10MiniMax M3minimax/minimax-m34.99%1.2×211%$0.19
10 runs

Provider pin: minimax/fp8

Arithmetic mean: 5.50%

  1. R14.00%
  2. R28.00%
  3. R35.00%
  4. R43.00%
  5. R56.00%
  6. R65.00%
  7. R76.00%
  8. R82.00%
  9. R911.00%
  10. R105.00%

10 parse failures · 0 infrastructure failures

Bars use a fixed 0–100% scale. The vertical mark is 4% chance. A zero geometric score can mean one repetition scored zero; no epsilon is added.

Four ways to recognize one tune

Condition matrix

Absolute geometric scores. Select a condition cell to move the leaderboard to that slice.

ModelOverallOriginal · Notes onlyOriginal · Pitch + durationTransposed · Notes onlyTransposed · Pitch + duration
Gemini 3.6 Flash41.36%
GPT-5.6 Sol40.50%
Claude Opus 4.835.36%
Kimi K331.59%
Claude Sonnet 522.51%
GPT-5.6 Luna22.34%
Qwen 3.7 Max20.12%
GLM 5.214.98%
DeepSeek V4 Pro10.06%
MiniMax M34.99%

Darker cells indicate higher scores · every cell is labeled · chance is 4%

Why it matters

Intelligence should travel between senses.

Humans can look at a silent line of notation and experience the shape of a melody. MeloText-ID asks whether language models can perform an analogous representational crossing.

A note string and a melody concept begin in different regions of representational space. Recognition requires a mapping that preserves interval relationships, rhythmic contour, and identity as the key changes.

01

Structural reconstruction

Tokens become an ordered melodic contour.

02

Temporal binding

Duration codes add rhythmic structure.

03

Invariant identity

Transposition changes pitch, not the tune.

Research context

From score to controlled test

How the benchmark was built

  1. 01

    Choose

    25 highly familiar, nameable melodies with clear monophonic motifs.

  2. 02

    Extract

    The most recognizable passage, checked against score sources.

  3. 03

    Encode

    Notes-only and duration-prefixed strings in 32nd-note units.

  4. 04

    Transpose

    Move the key while preserving intervals, attacks and measures.

  5. 05

    Randomize

    25 titles per item, with every correct answer position balanced.

  6. 06

    Repeat

    Ten independently randomized 100-item runs, graded exactly.

One melody, two textual surfaces

Notes onlyC C G G A A G

Pitch + duration8C 8C 8G 8G 8A 8A 16G

Task, prompt and stimulus construction

Every stateless request contains one symbolic melody and all 25 candidate titles. Models receive no audio, browsing, tools, metadata, or cross-request context. `R` means rest; accidentals and octave markers are encoded explicitly.

Download the frozen prompt ↓
Randomization, balance and grading

Candidate permutations are deterministic. Within each condition and repetition, every answer position is correct exactly once. Exact JSON choice IDs are graded automatically; malformed fully received answers remain incorrect.

Scoring and aggregation

Within each repetition, accuracy is arithmetic. Across ten repetitions, the score is the exact geometric mean. For a custom slice, selected cells are pooled inside each repetition before taking the geometric mean. No epsilon is added.

score = (A₁ × A₂ × … × A₁₀)^(1/10)

Model routing, costs and audit

Exact model/provider routes were pinned and fallback was disabled. The complete scored panel cost $39.111479. Hash, route, balance, scoring, export and cost checks passed.

Provenance and reproducibility

Version 0.1.3-pilot · finalized July 23, 2026 · source run live-20260722T122046Z-a0d2fe0276.

Experiment fingerprint: a0d2fe02767f66b48a6cce5a993698b44786efb78aab917b27f73225a2ef385c

The recognition set

25 melodies people remember.

The pilot favors recognizable motifs and reproducible monophonic encodings. It is not a balanced survey of global musical culture.

  1. 011812 Overture
  2. 02Ave Maria (Schubert)
  3. 03Beethoven's Fifth Symphony
  4. 04Ode to Joy (Beethoven's Ninth Symphony)
  5. 05Bella Ciao
  6. 06The Blue Danube
  7. 07Canon in D (Pachelbel)
  8. 08Clair de Lune
  9. 09Eine kleine Nachtmusik
  10. 10Fur Elise
  11. 11Hallelujah Chorus
  12. 12Happy Birthday to You
  13. 13In the Hall of the Mountain King
  14. 14The Internationale
  15. 15La donna e mobile
  16. 16March of the Volunteers (Chinese National Anthem)
  17. 17Mo Li Hua (Jasmine Flower)
  18. 18O Sole Mio
  19. 19Pomp and Circumstance March No. 1
  20. 20Queen of the Night Aria
  21. 21Rondo Alla Turca (Turkish March)
  22. 22The Star-Spangled Banner
  23. 23Toreador Song
  24. 24Two Tigers (Frere Jacques melody)
  25. 25Spring (Vivaldi, The Four Seasons)

Read the result precisely

Provenance & limitations

What this benchmark shows

Behavior consistent with cross-format musical reconstruction: parsing symbolic notation, preserving relational structure, and mapping it to a familiar semantic identity.

Cultural and training-data confounds

The small set is weighted toward Western classical music. Cultural familiarity and musical training matter, and familiar note sequences may have appeared in pretraining data.

What the representation omits

Monophonic text removes harmony, timbre, dynamics, articulation and performance. Closed-set recognition is not melody generation or general musical ability.

Why is Grok absent?

Grok 4.5 was stopped before completion after network instability and excluded entirely. No partial Grok result contributes to any rank or aggregate.

Official versus exploratory

The all-condition leaderboard is the official result. Every subset ranking is a transparent diagnostic view and should not replace the preregistered score.

Human comparison

This pilot contains no human baseline and cannot establish subjective auditory imagery, literal synesthesia, consciousness, or the internal method used by a model.

Open aggregate evidence

Inspect the benchmark.

Public files contain aggregate scores, protocol and integrity commitments—never raw answers or an answer key.

Independent research

Created by
Alejandro Zarzuelo.

MeloText-ID was conceived and created to investigate how AI systems connect symbolic structure, imagined sensory form, and semantic identity.

Visit alejandrozarzuelo.com ↗