Llama 4 Scout is the lightweight, single-GPU Llama 4 with a record 10M-token context — the most self-hostable open multimodal model for cost-constrained teams.
Representative public scores (approximate, higher is better) for relative comparison. Check the provider for the latest official results.
Free to self-host (Llama 4 community licence). 17B active / 109B total MoE; fits a single high-end GPU.
Access Llama 4 Scout