蚂蚁 inclusionAI 发布 LLaDA2.2-mini 智能体扩散语言模型

蚂蚁 inclusionAI:HuggingFace 新模型·2026-09-05 13:16·23小时前
AI 导读

蚂蚁 inclusionAI 在 Hugging Face 发布 LLaDA2.2-mini,这是 LLaDA2 系列的轻量级智能体扩散语言模型,总参数 16B、推理仅激活 1.4B。

蚂蚁 inclusionAI:HuggingFace 新模型
53AI 编辑部评分,满分 100

蚂蚁 inclusionAI 发布 LLaDA2.2-mini 智能体扩散语言模型

2026-09-05 13:16· 23小时前
AI 导读

蚂蚁 inclusionAI 在 Hugging Face 发布 LLaDA2.2-mini,这是 LLaDA2 系列的轻量级智能体扩散语言模型,总参数 16B、推理仅激活 1.4B。

LLaDA2.2-mini is the lightweight variant of the agentic diffusion language model in the LLaDA2 series. Built upon the LLaDA2.0-mini architecture, it inherits the core innovations of the LLaDA2.2 series — Levenshtein Editing (introducing DELETE and INSERT control tokens) — enabling long-context tool calling, multi-turn interaction, and robust error correction, while maintaining a smaller parameter footprint and lower inference cost. For more details, please refer to our technical report.


📊 Benchmarks

The following table compares LLaDA2.0-mini, LLaDA2.1-mini, and LLaDA2.2-mini across General and Agentic capabilities.

Category Benchmark LLaDA2.0-mini LLaDA2.1-mini LLaDA2.2-mini
General
Function Calling BFCL v4 25.05 28.44 47.68
BFCL v3 70.72 72.06 69.02
Math AIME 2026 37.71 40.37 35.05
OlympiadBench 67.70 64.30 61.11
Coding LiveCodeBench v6 27.70 28.80 28.14
MultiPL-E 67.46 64.16 65.26
Instruction Following IFBench 32.33 31.60 24.93
Multi-IF 60.58 58.43 57.03
Reasoning KOR-Bench 49.92 46.64 43.60
Knowledge GPQA-Diamond 47.76 48.36 44.41
Long Context LongBench v2 15.51 12.13 34.99
General Average 45.68 45.03 46.47
Agentic
Agent τ²-Bench - - 57.50
Claw-Eval - - 57.16
PinchBench - - 62.33
Agentic Average - - 59.00

🚀 Key Features

  • Efficient 128K Diffusion Infrastructure: LLaDA2.2-mini extends the context window to 128K and introduces the Block Routing mechanism, which restricts MoE expert activation at the diffusion block level, enabling efficient long-context agentic tasks.

  • Levenshtein Editing: Introduces DELETE and INSERT control tokens, enabling diffusion decoding to edit sequence structure, remove redundant content, and create insertion points during parallel generation.

  • Agentic Reinforcement Learning: Proposes Levenshtein Editing ELBO-based Block-level Policy Optimization (L-EBPO), leveraging agentic environment rewards to train Levenshtein editing and error correction capabilities in multi-turn tool-use scenarios.

  • Lightweight & Efficient: With a total of 16B parameters and only 1.4B activated during inference, it significantly reduces computational cost while maintaining strong capabilities.


📦 Model Variants

Model ID Description Hugging Face Link
inclusionAI/LLaDA2.2-flash Agentic MoE Diffusion Language Model (100B) with Levenshtein editing capabilities. 🤗 Model Card
inclusionAI/LLaDA2.2-mini Lightweight Agentic MoE Diffusion Language Model (16B) with Levenshtein editing capabilities. 🤗 Model Card

🔍 Model Overview

Key specifications of LLaDA2.2-mini:

  • Type: Mixture-of-Experts (MoE) Diffusion Language Model with Levenshtein Editing
  • Context Length: 128K tokens
  • Levenshtein Editing Control Tokens: DELETE, INSERT
  • Total Parameters (excl. Embedding): 16B
  • Layers: 20
  • Attention Heads: 16
  • KV Heads: 4
  • Experts: 256 (8 activated per token)
  • Positional Encoding: Rotary Position Embedding (RoPE)
  • Vocabulary Size: 157,184

🤗 Hugging Face Transformers Usage

Please ensure transformers>=5.2.0 and related dependencies are installed.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_path = "inclusionAI/LLaDA2.2-mini"

model = AutoModelForCausalLM.from_pretrained(
    model_path,
    trust_remote_code=True,
    device_map="auto",
)
model = model.to(torch.bfloat16)
model.eval()

tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)

prompt = "Calculate 1+5-28*0.5-200=?"
input_ids = tokenizer.apply_chat_template(
    [{"role": "user", "content": prompt}],
    add_generation_prompt=True,
    tokenize=True,
    return_tensors="pt",
).input_ids

generated_tokens = model.generate(
    inputs=input_ids,
    eos_early_stop=True,
    gen_length=512,
    block_length=32,
    threshold=0.5,
    editing_threshold=0.0,
    temperature=0.0,
)

generated_answer = tokenizer.decode(
    generated_tokens[0],
    skip_special_tokens=True,
)
print(generated_answer)

Best Practices

For optimal performance, we recommend the following configurations:

  1. Sampling Parameters: Use block_length=32, temperature=0.0, top_p=None, top_k=None as stable defaults.

  2. Denoising Threshold: Adjust threshold, editing_threshold, and max_post_steps based on the speed-quality trade-off for your use case. Lower thresholds can improve inference speed but may lead to repetitive or unstable outputs.

  3. Output Length: For most queries, an output length of 32768 tokens is recommended.

  4. Long-Context Agentic Tasks: For long-context tool calling and multi-turn agentic applications, we recommend using SGLang as the serving backend. Ensure the server configuration supports a 128K context window and the model's MoE diffusion inference requirements.


🤖 ModelScope

If you are in mainland China, we strongly recommend accessing our models via 🤖 ModelScope.


🌐 License

This project is licensed under the Apache License 2.0.


🤝 Contact & Collaboration

For any questions, collaboration opportunities, or feedback, please reach out to us via Hugging Face or submit an issue on our GitHub repository.

Join us in advancing open, efficient, and intelligent diffusion language models for agentic applications!


16B params

Collection including inclusionAI/LLaDA2.2-mini

来源:蚂蚁 inclusionAI:HuggingFace 新模型· huggingface.co