Show HN: Distill and serve small models with frontier quality for half the cost
Sentiment Mix
Geography
Expert Signals
SilenN
author • 1 mention
Hacker News
source • 1 mention
AI-Generated Claims
Generated from linked receipts; click sources for full context.
Today we are launching `wmo serve`, a tool to route repetitive tasks to distilled smaller models.Agent traces you already capture are opportunities to get signal on how to make your model cheaper, faster, better.
Supported by 1 story
We continuously improve - your specialized model through distillation from open source models - model routing to frontier + custom models - token compaction to remove noise and save tokensDemo: https://www.youtube.com/watch?v=2_m4Ze6mdkoPass in traces and an OpenRouter key, and wmo starts a local OpenAI-compatible endpoint to run with your model at a lower cost with equivalent quality.
Supported by 1 story
Tinker continually trains as new traces arrive.We also offer a hosted solution for anyone that just wants a frontier quality endpoint with self-improvement over time at a 40%+ lower cost.Sign...
Supported by 1 story
Related Events
AI Distillation Is a Standard Tool. Washington Is Deciding When It Becomes Model Theft - MLQ.ai
Uncategorized • 7/26/2026
Show HN: HART OS – an open-source AI OS built so frontier AI needs no datacenter
Open Source • 7/27/2026
Show HN: External KV Cache Offloading Cuts Long Horizon Inference Costs by 50%
Uncategorized • 7/26/2026
Show HN: Boffin – Staff-engineer layer for AI coding agents
Uncategorized • 7/27/2026
Show HN: Reproducibility Benchmark a Risk Quantitative Model
Research • 7/26/2026