AI Engineer LLM Systems
Прямой работодатель Atlantix ( www.atlantix.cc )
Опыт работы любой
🚀 AI Engineer — LLM Systems // Evals @ Ventora
Location: Remote 🌐
Employment Type: Full-time
Level: Mid-level / Senior
🌍 About Ventora
Ventora is an AI platform that helps people turn ideas into real products — from business planning and MVP creation to payments, marketing, and market validation.
The core platform is already built and in use. We’re now entering a stage of active growth and are looking for an AI-first engineer to help us make our AI systems smarter, more reliable, and easier to improve.
This is a hands-on product engineering role with real ownership. You’ll build LLM-powered features, improve AI pipelines, and create practical evaluation systems that show us what works and what needs to be fixed.
🎯 Your Mission
Help Ventora ship better AI systems — and make their quality measurable.
You’ll work across the product’s AI pipeline: business plan generation, app building, code generation, review, and marketing.
You’ll:
- Build and improve LLM-powered product features
- Develop multi-stage AI agents and workflows
- Integrate and test new models
- Improve prompts, context, routing, and fallbacks
- Create automated evals and regression tests
- Measure quality, cost, speed, and reliability
- Turn evaluation findings into product improvements
- Contribute to regular product engineering when needed
This is not a research-only or eval-only position. You’ll ship production code and see your work directly affect the product and its users.
🧩 What You’ll Work On
🤖 AI Systems & Product Features
Build and improve AI-powered features across Ventora’s product journey.
This may include multi-stage workflows such as plan → build → review, model integrations, structured outputs, tool calling, streaming, caching, prompt improvements, and new AI agents.
You’ll also contribute to production features, bug fixes, and pipeline development in the main codebase.
🧪 Evals & Quality
Help us understand whether each product change actually makes the AI better.
You’ll create practical evaluation systems for generated applications, business plans, marketing content, AI reliability, and release regressions. Depending on the task, you may use automated checks, browser tests, LLM judges, human-reviewed datasets, and side-by-side comparisons. When something fails, you’ll help find the cause, improve the prompt or pipeline, ship the fix, and add a regression test so the same issue does not return.
The goal is simple: catch problems before users do and make product decisions based on evidence.
🏆 What Success Looks Like
- New AI features and improvements reach production regularly
- AI quality becomes visible and measurable
- Regressions are caught before release
- Generated products become more functional and reliable
- Evaluation findings consistently turn into shipped fixes
- Model quality, cost, and latency stay under control
Over time, you’ll become a key technical partner for the founders, product team, and engineers shaping Ventora’s AI platform.
🤝 What You Bring
Required:
- 3+ years of experience in AI/ML, applied data science, or backend engineering
- Hands-on experience building production applications with LLMs
- Experience with model APIs such as OpenAI, Anthropic, or OpenRouter
- Understanding of prompts, agents, tool calling, structured outputs, and multi-stage AI workflows
- Ability to turn unclear questions like “Is this version better?” into practical, measurable tests
- Strong ownership, product thinking, and clear communication
- Regular use of AI coding tools such as Codex, Claude Code, Cursor, or similar
- Comfortable working in Linux environments and with Docker
You don’t need experience with every framework or evaluation platform. We care more about strong engineering fundamentals, curiosity, and the ability to learn and ship.
✨ Nice to Have
- Experience with LLM evals, datasets, scorers, or CI quality gates
- TypeScript or Node.js experience
- Browser automation, SQL, or production monitoring experience
- Familiarity with tools such as promptfoo, DeepEval, Ragas, Braintrust, LangSmith, or Langfuse
- Experience with agent frameworks, RAG evaluation, vision models, or AI red-teaming
- A background in mathematics, statistics, physics, machine learning, or another quantitative field
- Early-stage startup or side-project experience
🌱 Why Join Ventora
- Build real AI systems used in production
- Influence both the product and its technical foundation
- Own projects from idea to launch and measurement
- Work directly with founders, product, and engineering
- Experiment with new models and AI development workflows
- Help define quality standards for an AI-native platform
- Grow into a leading role as the product and team expand
If you enjoy building with LLMs, care about whether AI systems actually work, and want your work to have a direct product impact, we’d be glad to talk.
Join us and help build the next stage of Ventora!
