TinyAudio: Compact and Efficient Text-to-Audio Generation for Low-Resource Deployment

Junxi Liu1,2, Xiquan Li1,3, Wenhao Guan2, Yifan Duan1,2, Zhikang Niu1,2, Yanru Huo1, Ziyang Ma1,2,4, Xie Chen1,2,*
1X-LANCE Lab, Shanghai Jiao Tong University  ·  2Shanghai Innovation Institute
3SJTU Paris Elite Institute of Technology  ·  4Nanyang Technological University

*Corresponding author
AudioCaps quality versus generator size


TinyAudio is a compact flow-matching text-to-audio model that generates 44.1 kHz audio with only 87M inference-time parameters — over 90% fewer than representative billion-parameter pipelines — and 0.48 GB of peak GPU memory. Its core TA-DiT jointly models text and audio tokens in a single stream with layer-shared conditional modulation, while TA-CLAP provides audio-aligned text conditioning. TinyAudio-MF, trained with the Improved Mean Flows objective, reduces latent generation from 25 Euler steps to a single step, enabling real-time generation under a four-core CPU quota.


Overall Architecture


TA-DiT is TinyAudio's compact generative backbone. It removes two sources of repeated parameters in text-conditioned generative Transformers: separate modality-fusion modules and block-specific conditional modulation networks. Projected text tokens and noised audio latents are concatenated for joint self-attention, with independently indexed rotary position embeddings preserving the positional structure of each modality; only audio tokens predict latent velocity. Instead of a per-block AdaLN network, the shift–scale and gating MLPs are shared across all Transformer blocks, with lightweight block-specific embeddings preserving depth-dependent behavior. A single 18-block stream with hidden size 384 and 8 attention heads yields a 35M generator.


Single-stream TA-DiT architecture


Main Results


TinyAudio attains an AudioCaps FAD of 2.65 with 0.48 GB of peak GPU memory, more than 8× less than every compared system, and remains competitive in CU and PQ on TTA-Bench. TinyAudio-MF cuts sampling to a single function evaluation at a measurable quality trade-off.

Model Footprint / Efficiency AudioCaps TTA-Bench
Gen. Deploy. Mem. NFE RTF FAD↓ FD↓ KL↓ IS↑ CE↑ CU↑ PC↑ PQ↑ CLAP↑
AudioLDM-L-Full 739M953M3.89 GB200×21.190 4.3229.501.688.17 3.2755.1373.2195.8540.441
AudioLDM2-Large 718M1467M5.93 GB200×24.173 3.4525.531.657.98 3.4415.4402.9575.9840.415
Tango-Full 866M1295M5.28 GB200×21.880 2.6815.641.248.78 3.2645.1513.3605.9540.440
Tango-2-Full 866M1295M5.28 GB200×21.879 2.8715.931.2010.05 3.4685.1963.8005.8950.467
MeanAudio-L-Full 480M1032M4.79 GB250.057 2.5914.061.2112.46 3.3045.1512.9955.7430.469
EzAudio-XL 857M2120M8.61 GB200×21.896 3.6414.981.2911.38 3.3885.0553.6635.7190.446
GenAU-L-Full 1253M2031M18.54 GB200×20.937 2.0714.581.3610.43 3.3595.0963.5845.8020.468
TangoFlux 516M937M6.99 GB50×20.257 2.4120.651.2712.81 3.5395.0743.6735.7820.472
Resonate 470M1108M5.13 GB25×20.273 2.2715.531.3110.79 3.4515.3283.2326.0640.476
TinyAudio 35M87M0.48 GB25×20.078 2.6518.241.3210.44 3.3435.2622.9835.9960.429
TinyAudio-MF 35M87M0.48 GB10.010 2.8323.651.498.47 3.3505.2573.0435.9940.403

Bold and underline indicate best and second-best results. Mem. and RTF are measured on a single GPU in FP32 with batch size 1. TTA-Bench results use the Accuracy subset.