Avatar for Intelliswift Software
Intelliswift Software
Actively Hiring

Software Engineer/Research Engineer (AI Model Evaluation & Systems)

Posted: 4 weeks ago
Hires remotely in
Remote Work Policy

Remote only

Company Location
Newark
Visa Sponsorship

Not Available

RelocationAllowed
Skills
Python
PyTorch
Benchmarks
Deep Learning Frameworks
Validation Frameworks
Model Evaluation Workflows

About the job

Job Title: Software Engineer/Research Engineer (AI Model Evaluation & Systems)

Location: Remote (U.S.) PST Time Zone Preferred)

Duration: 12 Months (Potential Extension)

We are seeking a Software Engineer/Research Engineer to help advance next-generation AI systems by designing, building, and evaluating complex workflows used to assess model performance on real-world computer-based tasks. This role sits at the intersection of software engineering, machine learning, and AI evaluation.

You will work closely with researchers and engineers to develop benchmarks, test suites, automation frameworks, and evaluation pipelines that measure the capabilities of modern AI systems. The ideal candidate is highly technical, enjoys experimentation and debugging, and has a strong passion for building reliable, high-quality software.

Looking for a junior resource with 1-3 years max years of AI model training, or has extensive academic background in engineering AI.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Machine Learning, AI, or a related field
  • 1-3 years of experience in AI/ML engineering, research engineering, or software engineering
  • Strong programming skills in Python
  • Hands-on experience training and evaluating machine learning models
  • Experience with PyTorch and modern deep learning frameworks
  • Experience developing benchmarks, validation frameworks, or model evaluation workflows
  • Strong debugging, analytical, and problem-solving abilities
  • Demonstrated experience managing technical projects from implementation through validation and delivery
  • Familiarity with AI-assisted development tools such as Claude Code, Cursor, Codex, or GitHub Copilot

Preferred Qualifications

  • Experience with distributed training technologies such as DDP or FSDP
  • Knowledge of large language models (LLMs), generative AI, or AI agents
  • Experience with large-scale testing, developer tooling, or machine learning infrastructure
  • Exposure to cloud platforms such as AWS, Azure, or GCP
  • Experience with Docker, Kubernetes, CI/CD, or MLOps practices
  • Knowledge of Rust or TypeScript
  • Contributions to open-source projects, research publications, or personal technical projects

Responsibilities

  • Design, implement, and maintain evaluation frameworks for AI and machine learning systems
  • Develop benchmarks, test suites, and validation workflows used to measure model performance
  • Investigate model behavior, performance discrepancies, and system-level issues through data-driven experimentation
  • Build scalable Python-based tooling and infrastructure to support AI research and evaluation
  • Collaborate with researchers and cross-functional teams to improve model quality and performance
  • Document findings, methodologies, and technical decisions to ensure reproducibility
  • Contribute to deployment, monitoring, and operational excellence of AI systems
  • Drive engineering best practices around testing, reliability, and code quality

About the company

Funding

AMOUNT RAISED
$110M
FUNDED OVER
1 round
Round
S
$110000000
Seed - Nov 2024