Job Description
📋 Description LLM systems and prompt architecture: design prompts, outputs, and orchestration. Evaluation frameworks: build scalable AI-quality eval infra with tools. Model experimentation: run structured experiments for quality, cost, latency. AI observability and tooling: dashboards, debugging workflows. Matching and ML systems: improve match quality with lightweight models. Model serving and production inference: deploy, versioning, cost/latency optimization. 🎯 Requirements 5+ years of engineering experience building LLM-powered products. Strong prompt engineering and evaluation instincts. Experience building evaluation pipelines for LLM systems. Familiarity with vector databases and retrieval-augmented architectures. Experience with fine-tuning techniques such as SFT, DPO, RLHF. Experience with Python, SQLAlchemy, FastAPI, PostgreSQL. 🎁 Benefits Stock options with upside potential. Health and dental insurance. In-person work in NYC office.