News from Sapienza NLP

Sapienza NLP @ CLiC-it 2026

Sapienza NLP receives an Honorable Mention for Outstanding Contribution to Italian NLP Resources at CLiC-it 2026.

We are proud to announce that our paper Dromedario 3: Localizing the Tülu 3 Dataset into Italian has received an Honorable Mention for Outstanding Contribution to Italian NLP Resources at CLiC-it 2026 in Palermo! Congratulations to the whole team: Luca Gioffré, Marina Iuliana Aur, Francesco Ortame, Luca Moroni, Alberte Fernández-Castro, Elena Marafatto, and Roberto Navigli. Our contributions:

Dromedario 3: Localizing the Tülu 3 Dataset into Italian

by L. Gioffré, M. I. Aur, F. Ortame, L. Moroni, A. Fernández-Castro, E. Marafatto, and R. Navigli

Despite growing interest in supervised fine-tuning (SFT) datasets, nearly all existing resources are English-only or English-dominated, leaving Italian without a comparable resource. To bridge this gap, we introduce Dromedario 3, a large-scale Italian instruction-following dataset of approximately 630K translated instruction-response pairs derived from the Tülu 3 SFT mixture. Rather than translating the entire mixture, we first apply a multi-stage filtering pipeline. Specifically, we combine source-dataset filtering with automatic task and domain classification based on the Super-NaturalInstructions taxonomy to exclude language-bound tasks, while existing non-English instances are retained untranslated to preserve multilingual diversity. The selected instances are translated using TranslateGemma-27B and post-processed to remove faulty outputs. Manual validation confirms the reliability of both the automatic classification (accuracy above 87% for task labels and 95% for domain labels) and the machine translations (correctness above 95%). We further conduct a mixing study evaluating five SFT recipes with varying Italian-to-English compositions on parallel Italian and English benchmarks using two models, Llama 3.1 8B and Minerva 7B. The model with stronger Italian pretraining benefits substantially from Italian SFT data in generative settings, while recipes combining both languages offer the most robust trade-off between gains in Italian and preservation of English capabilities.