view article Article Illustrating Reinforcement Learning from Human Feedback (RLHF) +2 natolambert, LouisCastricato, lvwerra, Dahoas • Dec 9, 2022 • 434
Running 210 The ultimate guide to multi-harness RL 🔀 210 Train open models with RL inside real agent harnesses
deepseek-ai/DeepSeek-V4.1-Flash Image-Text-to-Text • 763B • Updated 8 days ago • 1.28M • • 4.27k
PII & De-Identification Collection Models for extracting PII entities and de-identifying clinical text, with support for HIPAA and GDPR compliance. • 366 items • Updated Jul 13 • 38
privacy-filter Collection OpenAI's privacy-filter fine0tuned models • 12 items • Updated Jul 13 • 12