Llama models will be decommissioned
Our upstream LLM provider is decommissioning all Llama models on 16 August 2026. After this date the following models will no longer be available through Helix inference:
llama-3.1-405b-reasoningllama-3.1-70b-versatilellama-3.1-8b-instantllama3-70b-8192llama3-8b-8192llama-guard-3-8b
What this means for you
Helix will automatically route inference to the next-best available model if a selected Llama model is unreachable. You may notice small differences in response style or latency. We recommend switching preferred models to qwen/qwen3.6-27b or another supported model before 16 August.
Timeline
| Date | Event |
|---|---|
| 14 Aug 2026 | Notice published and model alternatives confirmed. |
| 16 Aug 2026 | Llama models removed from Helix LLM routing. |
| 17 Aug 2026 | Incident review and routing validation report. |
Next steps
- Review any saved Helix preferences or custom BYOK keys that reference a Llama model.
- Run a test conversation with an alternative model to confirm output quality.
- Contact support@launchverse.app if you rely on a specific Llama behavior and need migration help.
