Granite 4.2 LLMs: How They're Built
17:14 · August 25, 2026 · Hugging Face Blog

Summary
Granite 4.2 comprises a family of dense decoder-only transformer models released in 3B, 8B, and 30B parameter sizes. The models are trained from scratch on approximately 15 trillion tokens through a five-phase curriculum that begins with broad web-scale data and progressively shifts toward curated sources before extending the context window to 512K tokens in the final phase. All three sizes share the same architectural backbone and training pipeline, with the larger variants receiving additional post-training to support agentic behavior.
Supervised fine-tuning mixes agentic trajectories (31.6 percent) with non-agentic instruction, coding, math, and multilingual data (68.4 percent), yielding roughly 7.2 million samples after normalization to OpenAI Chat format, LLM-based quality filtering, and both local and global deduplication. A second SFT stage for the 30B model upsamples agentic coding and software-engineering trajectories while retaining a replay buffer of earlier data. The resulting checkpoints then enter a multi-stage reinforcement-learning curriculum that applies Group Relative Policy Optimization asynchronously across separate environments.
The RL pipeline sequences verifiable-reward stages (RLVR and skill boosters) for all sizes, followed by an agentic block—covering software engineering, terminal interaction, and web search—for the 8B and 30B models only. Each stage runs as an independent GRPO job that warm-starts from the preceding checkpoint, using real-environment rollouts, group-relative advantages, and truncated importance sampling to accommodate asynchronous parameter refreshes. Final alignment employs RLHF with a modest KL penalty. Native tool-calling support is preserved throughout, allowing the models to emit OpenAI-compatible function calls that integrate directly with vLLM or SGLang endpoints.
The release also includes FP8 and GGUF quantization recipes, NeMo-RL training infrastructure, and benchmarks on reasoning, software-engineering, and agentic tasks. All variants are distributed under the Apache 2.0 license and incorporate Dutch-language coverage within the multilingual portion of the SFT mixture.
Why it matters
Provides production-grade details on training pipelines, RL methods, memory/latency optimizations, and deployment that ML engineers can directly apply or replicate in Dutch/EU settings.








