AI engineer · researcher
Stephanie Jarmak — AI engineer and researcher
I work on reliable AI systems: how models use context, tools, and structured knowledge to solve real problems, and how we evaluate whether those systems actually work. Currently an AI Engineer at Omni and a research affiliate with NASA Science Explorer (SciX).
Currently
- AI Engineer Omni · September 2026 – Present
- Research Affiliate NASA Science Explorer (SciX)
Selected work
All projects →-
SciX Agent
An agentic research assistant over the NASA SciX / ADS corpus, bridging AI agents with scholarly search infrastructure.
PythonAgentsMCPRetrieval
-
Gas City
An orchestration-builder SDK for multi-agent coding workflows. I'm a maintainer.
GoAgentsOrchestration
-
CodeScaleBench
A benchmark suite for evaluating how AI coding agents use external context-retrieval tools on realistic developer tasks in large, enterprise-scale codebases.
PythonEvaluationRetrievalDocker
-
EnterpriseBench
A benchmark for evaluating how well coding agents understand and navigate code across large, distributed enterprise codebases.
PythonEvaluationAgents
-
CodeProbe
Benchmarks AI coding agents against your own codebase by mining evaluation tasks from its git history, so the suite can't be contaminated by training data.
PythonEvaluationAgents
-
mem
Build and benchmark agentic memory using a multi-agent orchestrator's own work traces as the evaluation corpus, where every unit of work carries a real lifecycle outcome and a full trace.
TypeScriptPythonEvaluationAgent memory
Speaking
All talks →