About the Job
We’re searching for a hands-on AI engineer who can take modern LLMs from prototype to production, orchestrating multi-agent workflows using libraries such as LlamaIndex Workflows, LangGraph, and structured function calling. You should bring a solid foundation in classical ML (e.g., XGBoost) and deep learning with transformer-based models - especially LLaMA and Qwen-family models — along with experience in Retrieval-Augmented Generation (RAG) pipelines. A track record of building reliable, scalable systems is essential.
Key Responsibilities
- Proficiency in Python; experience with FastAPI and/or Java.
- Strong hands-on experience with ML models like XGBoost and DL frameworks like YOLO, OCR, and PyTorch.
- Practical knowledge of LLMs including fine-tuning, prompt engineering, and Retrieval-Augmented Generation (RAG)
- Familiarity with Hugging Face Transformers, PEFT, vector databases, and function-calling patterns.
- Experience with agent frameworks like LlamaIndex, LangGraph, or LangChain Agents. Skilled in deploying models using MLOps practices (Docker, CI/CD, experiment tracking, monitoring, rollback).
- Cloud experience with Azure ML / Azure Functions / AKS (preferred) or AWS SageMaker / Lambda.
Benefits of working here
- Best of Both Worlds: Enjoy the enthusiasm and learning curve of a startup combined with the deliveries and performance of an enterprise service provider.
- Flexible Working Hours: We offer a delivery-oriented approach with flexible working hours to help you maintain a healthy work-life balance.
- Limitless Growth Opportunities: The sky is not the limit when it comes to learning, growth, and sharing ideas. We encourage continuous learning and personal development.
- Flat Organizational Structure: We don't follow the typical corporate hierarchy ladder, fostering an open and collaborative work environment where everyone's voice is heard.
As part of our dedication to an inclusive and diverse workforce, TechChefz Digital is committed to Equal Employment Opportunity without regard to race, color, national origin, ethnicity, gender, protected veteran status, disability, sexual orientation, gender identity, or religion. If you need assistance, you may contact us at [email protected]
Qualifications
- Demonstrated ability to design and deploy ML/DL models for computer vision, NLP/GenAI, and tabular data.
- Experience orchestrating LLM agent behavior (state, memory, context) for complex task completion.
- Ability to build low-latency, secure inference APIs using FastAPI, including streaming support.
- Proven track record in cross-functional collaboration with product, data, and engineering teams.
- Bonus: Familiarity with Triton Inference Server, vLLM, websockets, or GPU cost optimization strategies.
Requirements
- Proficiency in Python; experience with FastAPI and/or Java.
- Strong hands-on experience with ML models like XGBoost and DL frameworks like YOLO, OCR, and PyTorch.
- Practical knowledge of LLMs including fine-tuning, prompt engineering, and Retrieval-Augmented Generation (RAG).
- Familiarity with Hugging Face Transformers, PEFT, vector databases, and function-calling patterns.
- Experience with agent frameworks like LlamaIndex, LangGraph, or LangChain Agents. Skilled in deploying models using MLOps practices (Docker, CI/CD, experiment tracking, monitoring, rollback).
- Cloud experience with Azure ML / Azure Functions / AKS (preferred) or AWS SageMaker / Lambda.
Location
Noida, Uttar Pradesh
Your new journey awaits!
Fill up a few details so that we can contact you regarding an opportunity.
Heads Up:
- Only pdf, doc and docx are accepted upto 5MB only.
- Applicants are appreciated to share their portfolio.