Skip to main content

Staff Research Engineer

AnyrayTel Aviv-Yafo, Tel Aviv District, IsraelNot specifiedFull-timeSeniority: Not specified

Posted yesterday · 0 applicants

Salary not listed for this role

Saving, applying or scoring takes a few seconds to set up your free account.

Willbi insight

The role in plain words

Must-have
  • Strong engineering fundamentals
  • Research instinct and ability to handle ambiguity
  • AI-native development experience using coding agents and LLM pipelines
  • Low-level knowledge of LLMs (tokenization, context windows, provider caching behaviors, streaming, tool calls, provider billing)
  • Staff-level experience carrying large/vague projects independently
Nice-to-have
  • Experience with inference infrastructure, LLM gateway/proxy, or model-serving efficiency
  • Experience building retrieval, embeddings, compression, or caching systems
  • Proficiency in TypeScript/Node and Python
  • Early-stage startup experience

Extracted from the job description · kept up to date automatically

Who this suits

Full job description

Original listing · kept for reference

Careers/Staff Research Engineer

Staff Research Engineer

Engineering Tel Aviv

• On-site / hybrid Full-time Staff level

We're hiring a staff research engineer to push Anyray's inference optimizations to the frontier — genuinely new work that lands on real enterprise traffic.

01The role

Anyray is a self-hosted, pre-model middleware that optimizes the cost and performance of every LLM request. That optimization lives across every layer of a request — input, context, tool output, files, dev commands, and caching.

The role is simple to state and hard to do well: make that optimization save more tokens without ever costing output quality or breaking a provider's cache. It's a research-heavy engineering role — you form theories from real traffic and ship the ones that prove out on live customer requests.

02Why this role

LLM optimization is a young field with no settled playbook, so the optimizations you invent here are genuinely new — and they ship into production, not a paper. And they matter: real enterprises run enormous volumes of LLM traffic and feel every wasted token, so the work you do lands on live customer requests and solves a problem they actually have. It's about as close to the frontier, and as close to the product, as engineering gets.

03What you'll do

• Find where the tokens go and cut them. You start from real enterprise traffic, form a theory about what's wasteful, and turn it into an optimization that ships to production and provably saves tokens.

• Prove the token savings. An idea that sounds good is worth nothing until the numbers back it. You'll build the evals and the token and quality measurement that tell a real win from a plausible one.

• Live in the caching trade-offs. Provider caches only help when the prefix is byte-for-byte stable, and a clever rewrite can easily burn more tokens than it saves. Knowing which is which is a lot of the job.

• Use AI to do the work of a team. We expect you to build coding agents, tooling and pipelines for yourself. The good ones end up as tools everyone here runs on.

• Set the bar. You take an optimization all the way, from a notebook to code serving live customer traffic, and the way you work rubs off on how the rest of us do.

04What we're looking for

• Strong engineering fundamentals, and the judgment to tell when something is actually correct, fast and safe rather than just green in CI.

• A real research instinct. You're fine with ambiguity, you form theories and drop them when the data says no, and you'd rather be right than clever.

• You already build AI-native. Coding agents and LLM pipelines are part of your day, and you've used them to do work that used to take a team.

• You know LLMs at a low level: tokenization, context windows, how caching actually behaves, streaming, tool calls, and how the various providers bill.

• Staff-level range. You've carried big, vague pieces of work on your own and shaped how a team builds.

05Nice to have

• You've worked on inference infra, an LLM gateway or proxy, or model-serving efficiency before.

• You've built retrieval, embeddings, compression, or caching systems.

• You're solid in TypeScript/Node, which is what the gateway and optimizer are written in, and comfortable in Python for research.

• You've been early at a company before and liked it.

06How to apply

If this sounds like you, email hi@anyray.ai. Tell us about an optimization you'd try first, or a time the data made you drop an idea you liked.

Apply

• TeamEngineering

• LevelStaff

• LocationTel Aviv

• Work modelOn-site / hybrid

• TypeFull-time

Name

Email

LinkedIn / GitHub

CV

• PDF or Word, ≤ 4 MB, optional

Why you

Prefer email? hi@anyray.ai

About Anyray
Company profile · coming soon

Employee reviews · coming soonMore roles at Anyray

Questions about this role

  • This listing did not state a salary. We only show pay when the employer publishes it.
Similar & related
Anyray
Posted yesterday · 0 applicants
See how you match