All roles

Open role

Research AI Engineer: Agentic Platform & Evaluation

Bay Area, CA · Seattle, WA

About AIDAChip

AIDAChip is the Alignment OS for AI-native silicon engineering teams — the first AI platform built for chip design teams, not just individuals. Chip design today is fragmented across disciplines, tools, and handoffs; we replace that with a multiplayer AI layer that keeps intent, knowledge, and execution coherent from specification to tapeout.

This role owns two tightly coupled areas at once: building our agent fleet and proving it works. You'll drive both the platform engineering (the agent machinery) and the research science (evaluation and publication) — owning the machinery that turns a domain expert's EDA agent design into a shipped production agent, while owning the rigorous testing that guarantees its performance.

Key responsibilities

  • Build the agent factory (harness, packaging, tool wiring) so a new role agent ships in weeks, not months
  • Design and ship orchestration agents (workflow coordination, verification orchestration, cross-discipline handoff)
  • Hold the fleet efficiency bar: minimize turns/task and agent token overhead, keep critical paths sub-second
  • Design the human-eval program with chip-design experts; build auto-evaluation that reaches parity with human judges
  • Research agent consistency: variance across repeated runs, drift across model swaps
  • Instrument the experience signals that gate the product: task completion, abandonment, HITL integrity
  • Pair daily with chip-design domain experts: they define what each agent must do and how it's judged; you ship it and measure it

You may be a good fit if you have

  • Shipped multi-agent / LLM-agent systems to production: tool-use harnesses, orchestration, durable long-running tasks
  • Deep cost/latency/reliability discipline for LLM systems at scale, with strong backend engineering
  • Rigor in eval methodology: inter-rater reliability, paired statistical testing, LLM-judge calibration
  • Production engineering ability: evaluation work ships code into CI gates, not notebooks

Strong candidates may also have experience with

  • The Python ecosystem for LLM and backend systems
  • A PhD (ML/NLP/IR or adjacent), an exceptional publication record, or an exceptional shipping-systems record

How we're different

  • Evidence over vibes: every behavior ships with a witness; claims verified before asserted; systems fail closed and loud
  • Measured, not narrated: a score not in the eval database didn't happen; benchmarks gate merges
  • Agent leverage: each engineer directs fleets of AI agents, shipping at a multiple of typing speed while holding the quality bar
  • Human-in-the-loop by design: agents propose; humans approve what matters
  • Greenfield ownership: a whole North-Star workstream, not a slice of a slice — on a production platform with real enterprise silicon customers, not a demo

Compensation

  • $175K – $250K • Plus Significant Equity

Apply

Send your application

Takes about a minute. We read every one.

Optional
Optional