Instructions to use TeichAI/Qwen3.8-27B-Fable-Distill with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TeichAI/Qwen3.8-27B-Fable-Distill with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="TeichAI/Qwen3.8-27B-Fable-Distill") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("TeichAI/Qwen3.8-27B-Fable-Distill") model = AutoModelForMultimodalLM.from_pretrained("TeichAI/Qwen3.8-27B-Fable-Distill", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use TeichAI/Qwen3.8-27B-Fable-Distill with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TeichAI/Qwen3.8-27B-Fable-Distill" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TeichAI/Qwen3.8-27B-Fable-Distill", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/TeichAI/Qwen3.8-27B-Fable-Distill
- SGLang
How to use TeichAI/Qwen3.8-27B-Fable-Distill with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "TeichAI/Qwen3.8-27B-Fable-Distill" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TeichAI/Qwen3.8-27B-Fable-Distill", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "TeichAI/Qwen3.8-27B-Fable-Distill" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TeichAI/Qwen3.8-27B-Fable-Distill", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Unsloth Desktop
- Docker Model Runner
How to use TeichAI/Qwen3.8-27B-Fable-Distill with Docker Model Runner:
docker model run hf.co/TeichAI/Qwen3.8-27B-Fable-Distill
Beta Testers Wanted for Qwen3.8-27B-Fable-Distill Deployment
Please do not use our community section to advertise your inference service.
Hi,
I'm using it since day 1 and it's really nice. (Q8) I usually write Kotlin and python codes with it. Compared to Qwen3.8 Flash Next, it has lesser tool usage, little lesser covarage of probabilities while thinking, but it doesn't make mistakes while coding, which makes it a perfect choice for making small adjustments in a project. Also, I'm Turkish and it can talk fluently, which is a big + for me. I wish you could do the same tweaks with pfeifferj--Qwen3.8-Flash-Next-GSQ-RCO-GGUF, because it's giving almost the same speed with better thinking in my setup. Also, adding some psychology datasets would be nice for chatting.
Thanks for everything :)
Hi, thank you for the feedback! Our goal was a light sft session to transfer agentic capability, I can get some more psychology, philosophy, and creative writing data for the future tunes. For that larger next model, we won’t be doing a tune since ~27B is our hardware limit for tuning. Stay tuned for Qwen4 though!