Back to search

Software Engineer (AI Platform Engineering)

XChange Software Inc

Hudson Signals
Prepare with AI feedback
Real-time coaching per answer
Apply to XChange Software IncOpens the original listing
50 hits10 practices6 interviews
Manhattan, New YorkContractAI/MLCloud Infrastructure

Software Engineer, AI Platform Engineering

12+ Months contract

This is a hybrid role in NYC , NY

Interview process:

1st round: 1 hour video interview

2nd round: In-person interview in NY office

The Agent Platform team would like to see candidates with Vue.js + Python experience

Senior Software Engineer, AI Platform Engineering

As a Senior Software Engineer, you will focus on the reliability and resilience of mission-critical AI platforms.

You will design systems that withstand provider and dependency failures, establish observability and service-level objectives, and keep the platform current as the AI ecosystem evolves.

The team’s scope includes Kubernetes-based platform-as-a-service frameworks, model-agnostic AI/LLM gateways, hybrid networking, observability, and self-service developer tooling.

We are looking for a collaborative, self-motivated engineer who is comfortable with ambiguity, takes ownership, and enjoys solving complex infrastructure challenges.

Legal AI is one of the most exciting and fast-moving areas in technology today.

If you are interested in building the foundational platforms that power the next generation of AI products, we'd love to hear from you.

What You’ll Do

  • Improve platform reliability and resilience by designing for failure, defining and meeting SLOs, leading incident response, and reducing operational toil.

  • Design, build, and operate Kubernetes-based PaaS frameworks, AI/LLM gateways, APIs, and self-service tools for AI applications across Bloomberg.

  • Develop model-agnostic gateway capabilities for providers such as OpenAI, Anthropic, Gemini, and AWS Bedrock, including routing, fallback, retries, rate limiting, and cost controls.

  • Build observability systems covering metrics, logs, traces, dashboards, and alerting to detect and resolve issues before they affect clients.

  • Develop networking solutions that connect applications across public-cloud and on-premises environments.

  • Provision and manage cloud infrastructure using Terraform and modern software engineering practices.

  • Keep platforms secure and current through dependency patching, runtime upgrades, migrations, and provider-integration updates.

  • Create frameworks, templates, and workflows that improve developer productivity and reduce operational overhead.

  • Evaluate emerging AI technologies and adapt the platform to support new development patterns and use cases.

What You’ll Bring

  • 6+ years of professional software engineering experience.

  • Strong Python skills and experience developing production-grade backend services and APIs; Java experience is a plus.

  • Experience designing and operating distributed systems in public-cloud environments, with a strong understanding of failure modes and resilient design patterns.

  • Hands-on AWS experience, including services such as EC2, S3, IAM, and container-based workloads.

  • Experience with Infrastructure as Code, preferably Terraform.

  • Experience with production operations, including metrics, logging, tracing, alerting, SLOs, and incident response.

  • Strong knowledge of software architecture, databases, networking, cloud infrastructure, and modern application development.

  • A degree in computer science, engineering, or a related field, or equivalent practical experience.

Preferred Qualifications

  • Experience building or operating API gateways, LLM gateways, or similar proxy layers with routing, fallback, rate limiting, caching, and cost tracking.

  • Experience with Open Telemetry, Prometheus, Grafana, Datadog, or similar observability tools.

  • Experience with Kubernetes and autoscaling technologies, preferably Amazon EKS and Karpenter.

  • Knowledge of AWS networking and security, including VPC, Direct Connect, IAM, and cloud security controls.

  • Experience with chaos engineering, load and failure testing, capacity planning, or disaster recovery.

  • Experience developing AI-powered applications, agent-based systems, model-inference services, or AI serving platforms.

  • Familiarity with AI development tools such as Claude Code, Cursor, or GitHub Copilot.

  • Working knowledge of machine learning concepts and the ML development lifecycle. Experience with SageMaker, Bedrock, PyTorch, TensorFlow, or scikit-learn is a plus.

  • The ability to learn quickly and independently lead large technical initiatives from concept through production.

AI enhanced job description

Source: jsearchRecruiter: Rekroot
Posted: 8/17/2026
70b6c9ac-d36c-4629-91a8-ef0c0b101ec1