Accordion with chirping birds in the background
Text prompt: “Accordion with chirping birds in the background.” Audio span: 4.4s–6.2s. Visual prompt: the accordion mask.
*Project lead
Multimodal representation learning has been shifting from two-tower architectures to LLM-based embedders. Existing models still treat text as the main way a user specifies intent, and they rarely cover video and audio together. OmniUE is an omni-interactive universal embedder. It learns a unified embedding for arbitrary combinations of text, video, and audio, and it conditions that embedding on user prompts given as text, visual regions of interest, or audio spans.
Visual and audio segmenters turn those prompts into entity- and span-level features. Learnable tokens read from intermediate layers of an omni-LLM are then aggregated into a single user-conditioned embedding. To test this setting we introduce OmniCHOIR, a compositional text-video-audio-to-audio retrieval benchmark with unimodal and multimodal interaction prompts.
Average relative gain over the strongest comparable baselines. Absolute scores are in the tables below.
OmniUE is built on Qwen2.5-Omni (3B and 7B). A query stream encodes the holistic input: text, silent video, and audio. A second stream encodes the interaction. SAM 3 segments visual prompts (points, boxes, or masks). SAM-Audio separates the sound indicated by a temporal span, text, or video. Connectors with convolution and self-attention map those feature maps into the LLM token space.
Four learnable tokens are appended to the sequence. Their hidden states are taken every four layers and fused with learned layer-token weights, then mean-pooled into one embedding. Training uses InfoNCE with GradCache on about 3.5M public multimodal samples. Visual and audio prompts are not in the training data: those interaction abilities emerge at test time, while text instructions are the training-time interaction.
Existing embedding benchmarks score text instructions on video or audio, or visual prompts on images. OmniCHOIR evaluates text, visual masks, and audio spans, alone and together, on text-video-audio-to-audio retrieval. Given a video, its mixed audio, and a prompt, the model retrieves the target sound together with its background from 16 candidates (Recall@1).
We start from held-out SAM-Audio-Bench clips, separate the target with SAM-Audio-Large, and use Qwen3-Omni-30B to propose a background category and hard negatives from ESC-50. A mixer injects the background at a random time and level (−15 dB to −5 dB). Each item has one reference, five background replacements, five same-category target swaps, and five different-category target swaps. After filtering, the benchmark has 479 ten-second clips.
Each clip is a query. The task is to pick the reference audio that matches the highlighted source and the background, not a same-class sound or a swapped background.
Text prompt: “Accordion with chirping birds in the background.” Audio span: 4.4s–6.2s. Visual prompt: the accordion mask.
Text prompt: “Female speaking with sea waves in the background.” Audio span: 5.0s–7.0s. Visual prompt: the speaker mask.
Best score in each column is highlighted. Rows tagged Ours are OmniUE.
| Model | CLS (5) | QA (5) | RET (5) | MRET (3) | Overall (18) |
|---|---|---|---|---|---|
| Text + image + video | |||||
| VLM2Vec-v2 (Qwen2-VL-2B) | 39.3 | 34.3 | 28.8 | 38.5 | 34.9 |
| UME-R1 (Qwen2-VL-2B) | 44.3 | 51.2 | 32.9 | 39.7 | 42.2 |
| UniME-V2 (LLaVA-OneVision-7B) | 37.2 | 50.6 | 28.9 | 39.6 | 39.0 |
| UME-R1 (Qwen2-VL-7B) | 48.6 | 60.7 | 38.2 | 39.3 | 47.5 |
| CAFe (LLaVA-OneVision-7B) | 35.8 | 58.7 | 34.4 | 39.5 | 42.4 |
| Text + image + video + audio | |||||
| Omni-Embed-Nemotron (3B) | 40.5 | 44.3 | 32.7 | 25.6 | 36.9 |
| LCO-Emb (3B) | 42.9 | 56.7 | 30.0 | 43.8 | 43.3 |
| e5-omni (3B) | 40.2 | 48.5 | 33.2 | 40.7 | 40.6 |
| LCO-Emb (7B) | 39.3 | 57.6 | 24.8 | 26.5 | 38.2 |
| e5-omni (7B) | 46.6 | 52.9 | 36.7 | 34.2 | 43.5 |
| WAVE (7B) | 50.7 | 45.9 | 34.7 | 39.5 | 43.1 |
| OmniUE-3B | 51.1 | 57.8 | 36.9 | 47.7 | 48.4 |
| OmniUE-7B | 57.8 | 62.9 | 40.9 | 41.3 | 51.8 |
Gains of OmniUE-3B / 7B over the best baseline of the same size: overall +5.1 / +4.3.
| Model | Clf | M.Clf | PC | Rrnk | Clust | A.Rtrvl | X.Rtrvl | Zero Clf. | Overall |
|---|---|---|---|---|---|---|---|---|---|
| MS CLAP 2023 | 45.0 | 5.8 | 53.6 | 75.4 | 15.2 | 87.3 | 9.4 | 12.6 | 31.1 |
| LAION CLAP | 51.7 | 2.3 | 51.9 | 66.8 | 6.6 | 93.2 | 9.8 | 14.9 | 32.2 |
| LCO-Emb (3B) | 56.4 | 41.6 | 66.7 | 75.4 | 1.3 | 67.7 | 50.3 | 62.2 | 50.7 |
| Qwen2-Audio-7B | 62.7 | 10.7 | 56.9 | 80.8 | 12.7 | 33.9 | 1.6 | 12.4 | 33.7 |
| LCO-Emb (7B) | 58.0 | 45.7 | 67.3 | 78.7 | 1.7 | 78.2 | 50.3 | 64.5 | 52.2 |
| OmniUE-3B | 57.3 | 35.3 | 67.5 | 88.4 | 1.7 | 81.1 | 46.5 | 63.5 | 50.5 |
| OmniUE-7B | 59.1 | 49.9 | 68.8 | 88.2 | 2.7 | 93.3 | 47.7 | 67.1 | 53.4 |
OmniUE-7B is best overall (53.4). Against LCO-Emb-7B it leads 7 of 8 meta-tasks. Clustering stays higher for Qwen2-Audio and the CLAP models, and cross-modal retrieval stays higher for LCO-Emb.
| Model | RefCOCO+ | RefCOCOg | VisualGenome | COCO-Stuff | ADE20K | Overall |
|---|---|---|---|---|---|---|
| VLM2Vec-2B | 24.5 | 29.5 | 22.3 | 19.4 | 24.6 | 24.1 |
| VLM2Vec-7B | 23.2 | 29.1 | 14.7 | 25.0 | 22.4 | 22.9 |
| MMRet-7B | 27.1 | 21.8 | 15.2 | 26.0 | 22.6 | 22.5 |
| UniME-7B | 31.4 | 32.8 | 19.0 | 25.3 | 23.0 | 26.3 |
| VIRTUE-2B | 28.8 | 42.4 | 24.4 | 29.9 | 27.5 | 30.6 |
| VIRTUE-7B | 33.0 | 35.3 | 19.6 | 27.1 | 23.8 | 27.8 |
| OmniUE-3B | 64.4 | 57.5 | 36.8 | 57.0 | 41.0 | 51.3 |
| OmniUE-7B | 65.3 | 66.2 | 35.5 | 61.6 | 47.5 | 55.2 |
Overall gain over the best 2B / 7B baseline: +20.7 / +27.4 points. The task is region-level caption retrieval from an image, a text instruction, and a box.
| Model | Text | Span | Mask | Text+Span | Text+Mask | Span+Mask | All three |
|---|---|---|---|---|---|---|---|
| LAION CLAP | 16.1 | 1.7 / 12.7 | — | 12.7 | — | — | — |
| ImageBind | 16.8 | 3.3 / 13.5 | — | 18.1 | — | — | — |
| LCO-Emb-3B | 19.6 | 18.7 | — | 20.5 | — | — | — |
| LCO-Emb-7B | 22.3 | 22.4 | — | 23.8 | — | — | — |
| WAVE-7B | 20.0 | 23.1 / 23.4 | — | 24.5 / 24.7 | — | — | — |
| OmniUE-3B | 23.6 | 25.3 | 25.6 | 23.6 | 25.7 | 25.7 | 26.4 |
| OmniUE-7B | 25.5 | 26.3 | 26.7 | 27.4 | 26.7 | 27.6 | 28.4 |
Recall@1. For baselines that cannot take a span, the pair is “span written as text / audio cropped to the span.” Masks and spans each beat text-only prompting, and using all three prompts is best (28.4).
@inproceedings{wang2026omniue,
title = {Omni-Interactive Universal Embedder},
author = {Wang, Wei-Yao and Tateishi, Kazuya and Cui, Shuyang and
Simon, Christian and Shibuya, Takashi and Takahashi, Shusuke and
Mitsufuji, Yuki},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026}
}