We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

ML Ops Engineer

Bana Solutions
United States, Virginia, Chantilly
Aug 04, 2026

OVERVIEW:

We are seeking an ML Ops Engineer to own the machine-learning lifecycle in production. You will be responsible for getting the five detection models from trained artifact to live, low-latency serving, then keeping them healthy -monitored, versioned, and retrained. Your product is the models running well in production, not the data pipeline underneath them.

GENERAL DUTIES:


  • Model release management in MLflow - versioning, aliasing, promotion and rollback, champion/challenger across the five models.
  • Serving models for real-time inference - package and optimize PyTorch models, run them in the low-latency inference workers, hold the <60s SLO; batching, CPU/GPU tradeoffs, inference correctness.
  • Model and prediction monitoring with Evidently - data, concept, and prediction drift; performance decay; alerting - and closing the loop back to retraining.
  • Automated retraining / continuous training - Airflow pipelines that retrain (including GPU training on EKS), validate against gates, and promote new model versions safely.
  • Training/serving consistency - manage the Feast online/offline boundary to prevent training-serving skew.
  • Reproducibility and governance - experiment tracking, model lineage/provenance, and model cards / approval gates for federal AI accountability.

REQUIRED QUALIFICATIONS:


  • Owned the full production ML lifecycle - trained artifact to live serving to be monitored/retrained. Not model-building only, and not data-pipeline-building only.
  • Model registry and experiment tracking - MLflow or equivalent (SageMaker, Weights & Biases, Vertex): versioning, promotion, rollback, lineage.
  • Model serving for real-time/low-latency inference - embedded serving or a model server (TorchServe, Triton, KServe, Seldon, BentoML): model loading, optimization, latency debugging.
  • Model and data drift monitoring - Evidently or equivalent; defining model-quality metrics and acting on decay.
  • Automated retraining / CT pipelines and model CI/CD - validation gates, champion/challenger, shadow or canary rollouts for models.
  • PyTorch (or TensorFlow) in production - packaging, optimizing (ONNX/quantization a plus), serving; debugging inference correctness and latency.
  • Feature store consumption (Feast or equivalent) with real focus on training/serving skew.
  • Kubernetes and Docker to package and deploy model workloads (Helm); Prometheus/Grafana for model and inference metrics.
  • Strong Python and solid software engineering (tests, reproducibility) - not notebook-only.


DESIRED QUALIFICATIONS:


  • The streaming pipeline you serve models into - Kafka + Bytewax (or Flink, Spark Streaming, Kafka Streams). You integrate with it; the data engineer owns it.
  • Apache Airflow used specifically for ML orchestration (training, promotion, drift jobs).
  • GPU training/serving on Kubernetes/EKS (CUDA/NVIDIA images).
  • OpenShift and/or air-gapped model deployment.
  • AWS GovCloud / FedRAMP / FIPS 140-2 / IL4-5, and federal AI governance - model cards, provenance, OSCAL, explainable scoring.
  • Graph ML, autoencoders, and anomaly detection (our detection approach); security/behavioral feature work.
  • Model artifacts in object storage (S3/MinIO); a warehouse (Redshift or equivalent) for offline evaluation data.


CLEARANCE:


  • Active U.S. Citizenship with the eligibility to gain a clearance
Applied = 0

(web-77cf7d65c7-jdxdg)