NeurIPS 2026

Omni-Interactive Universal Embedder

Wei-Yao Wang*, Kazuya Tateishi, Shuyang Cui, Christian Simon, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji
Sony Group Corporation  ·  Sony AI

*Project lead

OmniUE retrieves text, image, video, and audio from queries conditioned on text, visual regions, and audio spans.

OmniUE takes omni-interactive prompts and queries and retrieves targets in any modality: text, image, video, or audio. Prompts can be text, a visual region (click, box, or mask), an audio span, or any combination of them.

Abstract

One embedding space, any interaction

Multimodal representation learning has been shifting from two-tower architectures to LLM-based embedders. Existing models still treat text as the main way a user specifies intent, and they rarely cover video and audio together. OmniUE is an omni-interactive universal embedder. It learns a unified embedding for arbitrary combinations of text, video, and audio, and it conditions that embedding on user prompts given as text, visual regions of interest, or audio spans.

Visual and audio segmenters turn those prompts into entity- and span-level features. Learnable tokens read from intermediate layers of an omni-LLM are then aggregated into a single user-conditioned embedding. To test this setting we introduce OmniCHOIR, a compositional text-video-audio-to-audio retrieval benchmark with unimodal and multimodal interaction prompts.

+10.4%MMEB-v2-video
text-interactive video
+1.0%MAEB
text-interactive audio
+83.1%SCaR
visual-interactive
+21.9%OmniCHOIR
omni-interactive

Average relative gain over the strongest comparable baselines. Absolute scores are in the tables below.

Method

Query context plus a prompt stream

Text / video / audio query → SAM 3 + SAM-Audio prompts → Omni-LLM (Qwen2.5-Omni) → Learnable tokens → Context aggregation

OmniUE is built on Qwen2.5-Omni (3B and 7B). A query stream encodes the holistic input: text, silent video, and audio. A second stream encodes the interaction. SAM 3 segments visual prompts (points, boxes, or masks). SAM-Audio separates the sound indicated by a temporal span, text, or video. Connectors with convolution and self-attention map those feature maps into the LLM token space.

Four learnable tokens are appended to the sequence. Their hidden states are taken every four layers and fused with learned layer-token weights, then mean-pooled into one embedding. Training uses InfoNCE with GradCache on about 3.5M public multimodal samples. Visual and audio prompts are not in the training data: those interaction abilities emerge at test time, while text instructions are the training-time interaction.

OmniUE architecture: audio and vision segmenters, omni-LLM, learnable tokens, and context aggregation.
Segmenter streams (fire: trained, snowflake: frozen) inject user prompts. Context aggregation weights learnable tokens across layers to produce the embedding used for contrastive learning.
Benchmark

OmniCHOIR

Existing embedding benchmarks score text instructions on video or audio, or visual prompts on images. OmniCHOIR evaluates text, visual masks, and audio spans, alone and together, on text-video-audio-to-audio retrieval. Given a video, its mixed audio, and a prompt, the model retrieves the target sound together with its background from 16 candidates (Recall@1).

We start from held-out SAM-Audio-Bench clips, separate the target with SAM-Audio-Large, and use Qwen3-Omni-30B to propose a background category and hard negatives from ESC-50. A mixer injects the background at a random time and level (−15 dB to −5 dB). Each item has one reference, five background replacements, five same-category target swaps, and five different-category target swaps. After filtering, the benchmark has 479 ten-second clips.

OmniCHOIR collection pipeline from mixed video-audio to compositional audio candidates.
Collection pipeline. The reference mixes the separated target with a scene-plausible background. Negatives change the background, the target within its class, or both.

Examples

Each clip is a query. The task is to pick the reference audio that matches the highlighted source and the background, not a same-class sound or a swapped background.

Results

OmniUE vs. prior embedders

Best score in each column is highlighted. Rows tagged Ours are OmniUE.

MMEB-v2-video — text-interactive video

ModelCLS (5)QA (5)RET (5)MRET (3)Overall (18)
Text + image + video
VLM2Vec-v2 (Qwen2-VL-2B)39.334.328.838.534.9
UME-R1 (Qwen2-VL-2B)44.351.232.939.742.2
UniME-V2 (LLaVA-OneVision-7B)37.250.628.939.639.0
UME-R1 (Qwen2-VL-7B)48.660.738.239.347.5
CAFe (LLaVA-OneVision-7B)35.858.734.439.542.4
Text + image + video + audio
Omni-Embed-Nemotron (3B)40.544.332.725.636.9
LCO-Emb (3B)42.956.730.043.843.3
e5-omni (3B)40.248.533.240.740.6
LCO-Emb (7B)39.357.624.826.538.2
e5-omni (7B)46.652.936.734.243.5
WAVE (7B)50.745.934.739.543.1
OmniUE-3B51.157.836.947.748.4
OmniUE-7B57.862.940.941.351.8

Gains of OmniUE-3B / 7B over the best baseline of the same size: overall +5.1 / +4.3.

MAEB — text-interactive audio

ModelClfM.ClfPCRrnkClustA.RtrvlX.RtrvlZero Clf.Overall
MS CLAP 202345.05.853.675.415.287.39.412.631.1
LAION CLAP51.72.351.966.86.693.29.814.932.2
LCO-Emb (3B)56.441.666.775.41.367.750.362.250.7
Qwen2-Audio-7B62.710.756.980.812.733.91.612.433.7
LCO-Emb (7B)58.045.767.378.71.778.250.364.552.2
OmniUE-3B57.335.367.588.41.781.146.563.550.5
OmniUE-7B59.149.968.888.22.793.347.767.153.4

OmniUE-7B is best overall (53.4). Against LCO-Emb-7B it leads 7 of 8 meta-tasks. Clustering stays higher for Qwen2-Audio and the CLAP models, and cross-modal retrieval stays higher for LCO-Emb.

SCaR — visual-interactive text-image-to-text

ModelRefCOCO+RefCOCOgVisualGenomeCOCO-StuffADE20KOverall
VLM2Vec-2B24.529.522.319.424.624.1
VLM2Vec-7B23.229.114.725.022.422.9
MMRet-7B27.121.815.226.022.622.5
UniME-7B31.432.819.025.323.026.3
VIRTUE-2B28.842.424.429.927.530.6
VIRTUE-7B33.035.319.627.123.827.8
OmniUE-3B64.457.536.857.041.051.3
OmniUE-7B65.366.235.561.647.555.2

Overall gain over the best 2B / 7B baseline: +20.7 / +27.4 points. The task is region-level caption retrieval from an image, a text instruction, and a box.

OmniCHOIR — omni-interactive audio retrieval

ModelTextSpanMaskText+SpanText+MaskSpan+MaskAll three
LAION CLAP16.11.7 / 12.7—12.7———
ImageBind16.83.3 / 13.5—18.1———
LCO-Emb-3B19.618.7—20.5———
LCO-Emb-7B22.322.4—23.8———
WAVE-7B20.023.1 / 23.4—24.5 / 24.7———
OmniUE-3B23.625.325.623.625.725.726.4
OmniUE-7B25.526.326.727.426.727.628.4

Recall@1. For baselines that cannot take a span, the pair is “span written as text / audio cropped to the span.” Masks and spans each beat text-only prompting, and using all three prompts is best (28.4).

Cite this work

BibTeX

@inproceedings{wang2026omniue,
  title     = {Omni-Interactive Universal Embedder},
  author    = {Wang, Wei-Yao and Tateishi, Kazuya and Cui, Shuyang and
               Simon, Christian and Shibuya, Takashi and Takahashi, Shusuke and
               Mitsufuji, Yuki},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}