# Highlights - August 2026

By
Sebastien Frenck
Published 2026-08-01
  • Added inference-kimi-k3 (Kimi K3) to Sandbox Models for performance, stability, and cost validation; promotion to the Active Models catalog is pending successful evaluation. Hugging Face reference: moonshotai/Kimi-K3.

  • Added inference-minimax-h3 (MiniMax H3) to Sandbox Models for performance, stability, and cost validation; promotion to the Active Models catalog is pending successful evaluation. Hugging Face reference: MiniMaxAI/MiniMax-H3.

  • Added Document conversion with Docling, a managed document conversion API on the AI Model as a Service platform that turns PDFs, Office files, images and more into structured output such as Markdown, JSON and HTML using inference-qwen3-vl-235b.

  • Added inference-deepseek-v4-flash (DeepSeek V4 Flash) to Active Models on the Phoeniqs Model Service. Hugging Face reference: deepseek-ai/DeepSeek-V4-Flash-0731. Reasoning model with a 1M-token context window.

  • Scheduled inference-apertus-70b (Apertus 70B) for decommissioning on 21.08.2026; see Scheduled for Retirement. Suggested replacement: inference-apertus-v15-70b (Apertus v1.5 70B) on Active Models.

  • Added IPv4 address documentation to How Phoeniqs Cloud Works, covering subscription assignment to Capacity Pools, public internet reachability, and allocation rules.

  • Updated GPU Infrastructure and How Phoeniqs Cloud Works: users can enable VM passthrough or container GPU when provisioning a GPU subscription in the portal.