bottlecapai/ThinkingCap-Qwen3.8-27B
Image-Text-to-Text β’ 28B β’ Updated β’ 609 β’ 110
Efficient ThinkingCap Qwen3.8, bf16 and FP8 for datacenters. Request access and contact us for enterprise use if happy.
Note π bf16 reference, vLLM/SGLang/Transformers, 56 GB. Runs on any bf16 GPU (Ampere or newer). Needs an 80 GB GPU (H100, A100): weights plus ~2 GB per 32k tokens of context. The accuracy and reasoning-length baseline for FP8 below and for the quantized builds (NVFP4, GGUF, MLX): https://huggingface.co/collections/bottlecapai/thinkingcap-qwen-38-quants-6ab534cf5aee57103f21390d
Note π± FP8 block-wise, vLLM, 31.2 GB. Native on Hopper and Blackwell. Needs a 40 GB+ GPU (weights plus ~2 GB per 32k tokens of context). Evals on H200: accuracy β0.3 to +1.0 pp per benchmark, mean tokens β12% to +13%.