Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked.
Senior Software/ML Engineer — LLM Inference & Systems Optimization Research (5–10 Years of Experience)
Job description
AfterQuery is seeking experienced Senior Software/ML Engineers with deep expertise in LLM inference systems and AI infrastructure to help evaluate frontier AI models. In this fully remote contract role, you will apply your systems engineering knowledge to design challenging evaluation scenarios, create reference solutions, and assess AI-generated responses involving large-scale model serving, GPU optimization, and inference performance.
This is a research and evaluation position rather than a traditional software engineering role. Instead of building production systems, you'll help define what high-quality technical reasoning looks like for advanced AI models tackling complex systems engineering problems.
Key Responsibilities
Design Technical Evaluation Tasks
- Create realistic problem sets focused on LLM inference optimization and systems engineering.
- Develop challenging evaluation scenarios based on real-world infrastructure problems.
- Produce reference solutions that demonstrate expert-level reasoning.
Evaluate AI Systems
- Review AI-generated responses for technical correctness and engineering quality.
- Assess solutions involving model-serving infrastructure, GPU optimization, and inference performance.
- Provide detailed technical feedback to improve AI reasoning capabilities.
Contribute Systems Expertise
- Apply practical knowledge of large-scale AI infrastructure and performance optimization.
- Help establish evaluation standards for frontier AI systems.
- Collaborate with researchers advancing AI model capabilities.
Required Qualifications
5–10 years of professional software engineering or ML systems experience.
Bachelor's degree or higher in:
- Computer Science
- Computer Engineering
- Electrical Engineering
- Related technical field
Strong hands-on experience in one or more of:
- GPU inference optimization
- CUDA kernel development
- Model-serving systems (vLLM, TensorRT-LLM, SGLang)
- Quantization
- Optimizer design
Experience at a recognized technology or AI company with ownership of production systems.
Strong written communication skills for authoring technical explanations and feedback.
Preferred Qualifications
Experience with:
- SGLang
- Mamba or Mamba2 architectures
- IBM Granite models
Open-source contributions to inference or model-serving frameworks.
Experience with:
- Distributed inference
- Tensor parallelism
- Expert parallelism
- Continuous batching
- KV-cache optimization
Compensation
- $100–$150 per hour
- Independent Contract
Work Arrangement
- Fully Remote
- Contract engagement
- Flexible, part-time schedule
- Work on your own schedule
You will be redirected to the company's website to complete your application.