Cloudflare Releases Open-Weight 'Clef' and 'Clef-Flash' Decision Models on Workers AI
Production AI routing and agent workflows often stall while waiting for generative LLMs to emit token-by-token text. By introducing a prefill-only classification head, enterprises can execute multi-label schema decisions at $0.09 per million input tokens with median latency dropping to double-digit milliseconds.