AI Application Engineer (Part-Time)
About this role
Client: medxprts.ai
Location: Remote
Type: Part-time, with potential to transition to Full-time
Schedule: U.S. timezone overlap required
Description
Medxprts.ai is building an AI-powered platform for the legal and healthcare space, using LLMs, agentic workflows, and automation to create production-grade applications.
We are seeking an AI Engineer - LLM Fine-Tuning: a hands-on engineer who has personally trained and fine-tuned open-weight models, built training infrastructure, engineered datasets from messy documents, and established rigorous evaluation and preference/feedback training pipelines. This role works closely with engineering teams to ship models and integrated features into production.
Responsibilities
- Design, implement, and run LLM fine-tuning experiments (LoRA/QLoRA and full SFT) on open-weight models (e.g., Llama, Mistral, Qwen) and ship trained models into product workflows.
- Build and maintain training infrastructure using PyTorch and Hugging Face tooling (Transformers, PEFT, TRL/Axolotl), including multi-GPU training orchestration (DeepSpeed/FSDP) on AWS or GCP.
- Engineer datasets from real-world unstructured sources (long PDFs, medical/legal records), performing deduplication, filtering, contamination checks, and train/eval splits.
- Create evaluation harnesses tailored to domain needs: held-out test sets, LLM-as-judge with human calibration, regression tests across model versions, and automated monitoring for model drift.
- Implement preference/feedback training workflows (DPO/RLHF-style or similar) to learn from expert corrections (doctor-in-the-loop), and integrate feedback loops into model retraining pipelines.
- Collaborate with backend/frontend engineers to integrate fine-tuned models into services, optimize inference latency/cost, and support production deployments.
- Participate in PR reviews, release processes, incident debugging, and continuous improvement of training and deployment tooling.
- Ensure secure, compliant handling of sensitive data (HIPAA-awareness is highly preferred) during dataset preparation and model training.
- Hands-on experience fine-tuning LLMs: personally trained or fine-tuned open-weight models using LoRA/QLoRA and full SFT; able to explain trade-offs and provide at least one shipped example.
- Training infrastructure experience: PyTorch + Hugging Face ecosystem (Transformers, PEFT, TRL/Axolotl), multi-GPU training knowledge (DeepSpeed or FSDP), and running training workloads on AWS or GCP.
- Dataset engineering expertise: built instruction/preference datasets from messy, unstructured documents; practical knowledge of deduplication, filtering, train/eval splits, and contamination prevention.
- Strong evaluation discipline: designed domain-specific evaluation harnesses beyond standard benchmarks, including human-calibrated judge setups and regression testing.
- Practical experience with preference/feedback learning methods (DPO, RLHF-style workflows, or equivalent) and integrating expert feedback into model updates.
- Solid software engineering fundamentals: production workflows (Git, PRs), debugging, testing, deployment experience, and maintainable code.
- Experience with APIs, databases, and service integration for model inference.
- Ability to work independently, learn quickly, and follow technical direction.
- Good English communication skills and availability to overlap with U.S. working hours.
Nice to Have
- Experience with long-context handling strategies for very large documents (retrieval-aware training, context extension, RAG for multi-thousand-page sources).
- Model deployment and inference optimization: quantization (GPTQ/AWQ), vLLM/TGI serving, batching/throughput tuning, latency and cost optimization.
- Familiarity with vector databases, advanced RAG pipelines, MCPs, or n8n-style workflow automation tools.
- Knowledge of Docker, CI/CD, Kubernetes/EKS, or serverless infrastructure.
- Prior experience in healthcare, legaltech, HIPAA-aware processes, or other regulated/data-sensitive environments.
- Public portfolio, GitHub, or examples of shipped LLM/agentic applications and fine-tuning projects.
- Fully remote role.
- Opportunity to work on real AI products in the legal and healthcare domain.
- High ownership and autonomy.
- Performance-based incentives and outcome-driven bonuses.
- Potential to grow into a long-term, full-time role.
Frequently Asked Questions
Is the salary disclosed for the AI Application Engineer (Part-Time) position at Workana?
Is the AI Application Engineer (Part-Time) job at Workana remote?
Is the AI Application Engineer (Part-Time) role at Workana full-time or part-time?
Which team or department does the AI Application Engineer (Part-Time) at Workana belong to?
How do I apply for the AI Application Engineer (Part-Time) position at Workana?
When was the AI Application Engineer (Part-Time) job at Workana posted?
You'll be redirected to Workana's official application page on workable.