Back to articles
Speech & Audio

YuE2 Plans the Music Before Rendering a Complete Song

3 min read

Introduction

Music generation has traditionally split into two camps. Symbolic systems expose melody, harmony, rhythm, and form, but often stop at a score or event sequence. Audio systems can produce a finished recording, yet their compositional decisions are mostly hidden and difficult to edit. YuE2 targets this gap by placing an explicit musical plan in front of audio realization: the model first decides what the song should be, then renders it.

Key points

  • A score comes first. YuE2 uses a single AR-NAR Mixture-of-Transformers to generate a readable score describing melody and harmony. That representation is expanded into semantic music tokens and ultimately rendered as full-song audio.
  • One model connects planning and production. Rather than treating symbolic composition and waveform generation as unrelated stages, the system aims to unify composition planning, semantic representation, and audio realization. The resulting songs include vocals and accompaniment.
  • The reported tests favor planning. In a same-checkpoint comparison, experts gave 49.3% of overall preferences to the symbolic-planning version, compared with 34.6% for the version without planning. The result suggests that an explicit intermediate representation can affect not only editability but also perceived musicality.
  • Strong reported benchmark results. On WildSongBench, YuE2 received 6.73 on SongBench Global Avg, above the public baselines evaluated in the report. Selecting the best of eight candidates raised the mean to 6.96. The authors also report that best-of-8 outputs were preferred to Suno v4.5 and were nearly balanced against Suno v5 in expert listening.
  • Additional models support the training pipeline. Recordings rarely come with precisely aligned scores. To address that problem, the project presents MERT2 for music representation learning and SheetSage2 for symbolic transcription. The report says MERT2 surpassed prior best results on 14 of 15 MARBLE metrics, while SheetSage2 led on 12 of 15 benchmark-metric pairs in the authors’ evaluation.

Why it matters—and what remains open

The most interesting aspect of YuE2 is not simply its ability to produce a song, but its attempt to make generation inspectable. A score can serve as an intermediate control surface for changing melody, chords, or structure before committing to audio. The project also describes zero-shot covers, score-based editing, and editing guided by an external language-model agent, pointing toward a more interactive type of music software.

The results should still be read in context. They come from the authors’ reported benchmarks and expert listening studies, while best-of-8 performance depends on selecting among multiple generations. That does not necessarily mean every individual sample reaches the same level. It also remains unclear how well symbolic planning handles intricate arrangements, lyrics, vocal expression, and users who do not want to work with musical notation. Even so, YuE2 presents a compelling design direction: combine the controllability of symbolic composition with the immediacy of complete audio generation.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles