Chief Tokenminmaxing Officer (CTO)
פורסם אתמול · 0 מועמדים
תיאור המשרה המלא
המשרה המקורית · נשמר לעיוןKadima LaProduction is a political movement built by people who have spent too much time on call.
After an internal audit found eight unnecessary tokens in every production call, all spent on basic manners, we’re expanding our leadership team.
We’re looking for a Chief Tokenminmaxing Officer to own model selection, inference costs, and token policy across the organization. Your mandate is simple: reduce paid intelligence without causing a level of degradation anyone can reliably reproduce.
What you’ll do
• Define model-routing policy. Frontier models are reserved exclusively for investor demos and campaign content. Production will use what’s free.
• Set hard token budgets for prompts, RAG pipelines, memory, and agentic workflows. Requests for additional context require executive approval.
• Replace unnecessary multi-agent architectures with one quantized 3B model and, where possible, a normal Python function.
• Continuously migrate to whichever provider is offering free credits that quarter. Migration costs belong to Engineering.
• Build evals proving that no statistically significant number of users noticed the decline in quality.
What we’re looking for
• 7+ years in backend, ML platform, or infrastructure engineering, with production experience in model routing, RAG, caching, and quantization.
• A proven track record of cutting inference costs without turning the savings project into a larger infrastructure bill.
• Healthy suspicion of any workflow that needs a 128K context window to reset a password.
• Comfortable sharing one Claude Pro account with the leadership team. Access is allocated by seniority.
This role reports directly to Adir Commitov, Chairman of Kadima LaProduction. Applications exceeding 512 tokens will be truncated automatically.
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.