Micro1
Infrastructure Engineer
Posted
2 weeks ago
Experience
4/6+ Years
Salary
$20 - $70/hour
Deadline
Closed
Job Summary
The Infrastructure Engineer designs, builds, and evaluates complex synthetic cloud and on-premises technical environments to train frontier AI models and autonomous agents. Core daily duties include engineering realistic networking, storage, and containerization scenarios; developing infrastructure as code (IaC) deployment variations; constructing strict quality assurance grading rubrics to test model troubleshooting and incident response logic; identifying system failure patterns; and collaborating with data science teams to optimize training data quality.
Foundational Capabilities & Prerequisites
- Cloud Architecture Depth: Expert-level experience navigating and managing AWS, Azure, or Google Cloud Platform (GCP) environments.
- Baseline Infrastructure Knowledge: Strong professional background in core networking, enterprise storage, and system virtualization technologies.
- Automation & IaC Fluency: High proficiency utilizing central infrastructure-as-code and automation tools such as Terraform, Ansible, or similar frameworks.
- Containers & Orchestration: Deep structural understanding of container systems and orchestration engines, specifically Kubernetes and Docker.
- Diagnostics & Resolution: Advanced troubleshooting, system monitoring, and incident response capabilities.
- Articulation Skills: Exceptional written and verbal communication skills, showing a clear ability to translate complex technical concepts effectively for diverse audiences.
- Remote Autonomy: Demonstrated capability to thrive and execute high-quality data workflows independently within a fully remote, collaborative, and fast-paced environment.
Preferred Technical Multipliers
- Industry Accreditations: Relevant cloud or infrastructure certifications (e.g., AWS Certified Solutions Architect, Azure Solutions Architect).
- Governance Frameworks: Practical experience working with robust security practices and international compliance standards (SOC2, ISO 27001).
- Distributed Alignment: Prior work experience supporting global or distributed technical teams in a fully remote context.
- AI Evaluation Exposure: Past experience training AI models, evaluating LLM behaviors, or working in simulation/data design roles.
Technical Workflow Architecture
As an Infrastructure Engineer at micro1, your daily contributions will balance complex systems architecture with machine learning data design:
- Synthetic Architecture Prototyping: Designing extensive, complex cloud models featuring intentional misconfigurations—such as broken subnets, conflicting IAM policies, or invalid Dockerfiles—to evaluate whether AI agents can safely diagnose and fix systemic bugs.
- Adversarial Incident Simulation: Drafting highly specific engineering scenarios and system prompts that mimic live infrastructure failures—such as sudden container crashes, network bottlenecks, or security compliance breaches—to see if the AI model drops context or loses logical reasoning.
- Code Infrastructure Validation: Writing, structuring, and maintaining faulty Terraform scripts or Ansible playbooks to test if an AI engine can accurately identify execution errors, refactor code, and implement security best practices.
- Performance Failure Diagnostics: Auditing automated chat logs and terminal commands generated by AI devops tools, identifying structural patterns in their reasoning errors, and adjusting the training data to fix those specific algorithmic weaknesses.
Key Responsibilities
1. Synthetic Infrastructure & Catalog Architecture
- Design, deploy, and manage synthetic, scalable cloud and on-premises infrastructure solutions to act as training sandboxes for AI agents.
- Automate complex provisioning, configuration, and deployment simulation processes using Infrastructure as Code (IaC) tools.
- Structure rich technical environments incorporating networking setups, storage systems, and virtualization layers to test real-world complexity.
2. Scenario Design & Incident Response Simulation
- Develop diverse, authentic user prompts and system deployment scenarios simulating varied infrastructure demands, performance bottlenecks, and security challenges.
- Engineer challenging edge-case situations, including sudden system outages, security failures, and compliance gaps to test model resiliency.
- Implement and maintain robust security practices and compliance standards (SOC2, ISO 27001) across simulated training data.
3. Rigorous Evaluation, QA & Rubric Construction
- Create comprehensive evaluation rubrics and conduct rigorous QA to assess AI-generated code, architecture layouts, and troubleshooting responses for technical accuracy.
- Monitor system and network performance indicators within models proactively, identifying failure patterns to drive iterative improvement.
- Continuously analyze model performance outputs, optimize resource cost structures within data models, and refine data and scenarios for scalable AI model training.
4. Cross-Functional Data Science Collaboration
- Collaborate closely with AI and data science engineering teams to translate your practical infrastructure expertise into actionable training datasets and insights.
- Create comprehensive documentation and deliver clear, effective written and verbal communication to stakeholders to drive simulation-based AI training best practices.
Core Competencies & Analytical Strengths
- IaC & Logic Verification: Applying deep automation knowledge to spot script syntax errors, improper access permissions, and cloud configuration weaknesses in automated model payloads.
- Incident Scenario Simulation: Stepping into the mindset of a site reliability engineer to generate realistic system alerts and adversarial prompts that expose AI reasoning gaps.
- Analytical System Diagnostics: Systematically sorting through model outputs to diagnose exactly where and why an AI tool miscalculated network subnets or ignored firewall security rules.
- Asynchronous Data Scale: Designing organized, clean, and massive system data templates that data scientists can easily feed directly into model fine-tuning runs.
Expected Outputs & Deliverables
- Pristine, multi-cloud synthetic infrastructure topologies delivered in highly structured text formats for machine learning ingestion.
- A comprehensive suite of targeted infrastructure prompts, incident response challenges, and edge-case configuration scenarios.
- Verifiable quality assurance evaluation logs accompanied by detailed model error-tracking matrices.
- Standardized, repeatable rubric guidelines for scoring the cloud and networking performance of autonomous AI agents.
Skills Required:
- Computer / Software / It / Data
Quick Actions
Share Vacancy