Make compatible with Transformers 5.17 (#3) 4120e63 poolside-eng Yu-Peng-Poolside commited on 19 days ago
fix(rope): YaRN attention_factor is (0.1*ln(factor)+1)*attn_factor, not the bare multiplier 607820a verified joerowell commited on Jul 14
Use chat_template.jinja as the single source: drop the {% include %} chat_template field from tokenizer_config.json 2ce6e7f verified joerowell commited on Jul 1
Note FP8 KV cache needs vLLM 0.22.0; drop scrambled-output workaround (vllm#42650) 88ad86d joerowell commited on Jun 5
Drop VLLM_USE_DEEP_GEMM=0 from vllm serve recipe (DeepGEMM is supported on Hopper and datacenter Blackwell) b17bd8d verified joerowell commited on May 20
Enable thinking by default in non-Hopper FP8-KV serve command e8b30a7 verified joerowell commited on Apr 29
Update non-Hopper FP8-KV serve command and link to vLLM recipes page 2ad0232 verified joerowell commited on Apr 29