Capmation Inc.
Mexico City / Global
Mexico City / Global
The Role
We are seeking a Solutions Architect to join our Engineering Team. This role combines deep hands-on engineering capability in Python, machine learning frameworks, and data pipelines with technical leadership in designing, developing, deploying, and optimizing production-grade machine learning models and solutions.
The ideal candidate is a senior technical leader who can define end-to-end ML solution architecture, lead model development, evaluation, deployment, monitoring, and production support, and ensure adherence to ML engineering and MLOps best practices while delivering reliable, scalable, and maintainable solutions in a fast-moving environment.
This position also requires strong collaboration and leadership skills. The Solutions Architect must work effectively with data scientists, engineers, product stakeholders, clients, and delivery teams, while providing technical guidance, mentoring others, and driving high engineering quality. The ideal candidate should be proactive, pragmatic, and able to balance hands-on implementation with strategic technical decision-making.
Key Responsibilities
Solution Architecture: Lead the architecture of machine learning solutions end-to-end, from data ingestion and feature engineering to model serving and monitoring, ensuring scalability, security, and maintainability
Model Development: Design, develop, and optimize machine learning models using Python and ML frameworks, while leading evaluation and adoption of new tools, algorithms, and methodologies to improve company standards
Data Pipelines: Design and oversee reliable batch and streaming data pipelines and feature stores that feed training and inference, ensuring data quality, lineage, and reproducibility
Model Evaluation: Define evaluation strategies, metrics, validation approaches, and experiment tracking to ensure models meet accuracy, fairness, and business performance targets before release
Deployment & MLOps: Lead model packaging, deployment, and CI/CD/CT (continuous training) pipelines, including model registry, versioning, and automated promotion across environments
Monitoring & Production Support: Own production health of ML systems by monitoring model performance, data and concept drift, latency, and cost; lead incident triage, root-cause analysis, and retraining decisions
Best Practices: Establish and ensure adherence to ML engineering best practices, including code quality, testing, reproducibility, documentation, and responsible-AI and data-privacy requirements
Cross Functional Collaboration: Partner with business ops, stakeholders, and clients to translate business problems into ML solutions and act as the intermediary between business operations, data science, and engineering
Team Development: Provide technical guidance and mentor engineers and data scientists at all levels while also leading training sessions, design and code reviews, providing constructive feedback, and aligning technical standards across the team
Soft Skills
Business Acumen: Connect ML architecture and modeling decisions to business outcomes, anticipating impacts on cost, risk, and value, and providing decisions to maximize long-term value.
Accountability: Accountable for the technical and delivery success of ML solutions or projects, taking ownership of outcomes across teams and addressing issues proactively rather than reactively.
Communication: Communicate complex ML concepts, model behavior, and trade-offs clearly to both technical and non-technical audiences, aligning stakeholders, and enabling confident decision making.
Judgement: Demonstrate judgment by making high-impact decisions, balancing experimentation and short-term delivery with long-term sustainability, and escalating risks early.
Collaboration: Drive alignment across multiple teams and disciplines by acting as a unifying technical leader, resolving cross-team friction.
Curiosity: Maintain curiosity about the evolving ML landscape and emerging technologies, using that understanding to anticipate challenges, guide innovation, and continuously improve technical and delivery practices.
Required Qualifications
Experience: Over 6+ years of software or data engineering experience, including 3+ years in an architect or technical lead role and 4+ years delivering machine learning solutions to production, in the following:
Tech Stack
Languages & Core ML
Python (primary); SQL; Scala or Java a plus
NumPy, pandas, Polars
scikit-learn, XGBoost, LightGBM, CatBoost
Deep Learning Frameworks
PyTorch, TensorFlow / Keras
Hugging Face Transformers
Model optimization: ONNX, quantization, distillation
Data Engineering & Pipelines
Apache Spark / PySpark, Databricks
Orchestration: Apache Airflow, Prefect, Dagster, or Azure Data Factory
Streaming: Kafka, Azure Event Hubs, or AWS Kinesis
Data warehouses and lakehouses: Snowflake, Delta Lake, BigQuery
Data validation: Great Expectations, Pandera
MLOps & Model Lifecycle
Experiment tracking and model registry: MLflow, Weights & Biases
ML platforms: Azure Machine Learning, AWS SageMaker, Google Vertex AI
Feature stores: Feast, Databricks Feature Store, or equivalents
Pipeline frameworks: Kubeflow, Azure ML Pipelines, SageMaker Pipelines
Model Serving & APIs
FastAPI, Flask, or equivalent API frameworks
Serving: BentoML, KServe, TorchServe, Triton, or managed endpoints
Batch and real-time inference patterns, RESTful API design
Cloud & Infrastructure
Azure, AWS, or GCP cloud-native services
Containers and orchestration: Docker, Kubernetes
Terraform or Bicep for Infrastructure as Code
Architecture & Patterns
ML reference architectures (training, inference, feedback loops)
Event-driven and microservices-based ML systems
Reproducibility patterns: data and model versioning (DVC, Delta Lake)
DevOps
CI/CD/CT pipelines (Azure DevOps, GitHub Actions)
Git branching and pull request workflows
Testing & Quality
pytest, unit and integration testing for data and ML code
Model validation, bias/fairness testing, and explainability (SHAP, LIME)
Monitoring & Operations
Model and drift monitoring: Evidently AI, WhyLabs, Arize, or platform-native tools
Observability: Prometheus, Grafana, Application Insights, OpenTelemetry
Must have:
Proven experience designing, developing, deploying, and optimizing machine learning models in production using Python and ML frameworks (scikit-learn, PyTorch, TensorFlow, XGBoost).
Strong experience building and operating data pipelines for training and inference, including Spark / Databricks and workflow orchestration (Airflow or equivalent).
Hands-on MLOps experience: experiment tracking, model registry, versioning, and CI/CD/CT pipelines using MLflow and a major ML platform (Azure ML, SageMaker, or Vertex AI).
Track record leading model evaluation, deployment, monitoring (performance and drift), and production support for business-critical ML systems.
Experience deploying models as scalable services (batch and real-time) using containers, Kubernetes, and API frameworks such as FastAPI.
Track record leading solution architecture for enterprise-scale systems, including non-functional requirements, trade-off analysis, and architecture governance.
Demonstrated ability to define and enforce ML engineering best practices and to provide technical guidance and mentor engineering teams.
Preferred Qualifications
Experience with deep learning for NLP or computer vision, and with Hugging Face Transformers.
Experience with feature stores and real-time/streaming ML architectures (Kafka, Event Hubs).
Exposure to LLMs and generative AI, including fine-tuning, RAG, or LLMOps practices.
Familiarity with AI governance and risk frameworks (NIST AI RMF, ISO/IEC 42001), model explainability, and data privacy regulations.
Experience with Snowflake or lakehouse architectures (Delta Lake, Databricks Unity Catalog).
Consulting or client-facing delivery experience, including discovery and whiteboard sessions.
Cloud or ML certifications (e.g., Azure Data Scientist Associate, AWS Machine Learning Specialty / ML Engineer Associate, Google Professional ML Engineer).
Ciudad De México / Global
Mexico / Global
Distrito Federal / Global
Mexico / Global
Distrito Federal / Global
Mexico City / Global