Data Center System Software Architect, DGX Cloud

1 Month ago • 10 Years + • DevOps • Research & Development • $184,000 PA - $425,500 PA

Job Summary

Job Description

NVIDIA seeks a Data Center System Software Architect for its DGX Cloud team. Responsibilities include leading the architecture, design, and implementation of next-generation DGX cloud clusters using cutting-edge technologies. This full-stack role encompasses hardware architecture, workload orchestration, and application performance tuning. The ideal candidate possesses 10+ years of experience in system software, strong programming skills (C, C++, Go, Rust), expertise in distributed systems, and excellent communication skills. The role involves collaborating with various engineering teams across NVIDIA to ensure seamless software integration, from hardware to AI training applications. The architect will provide solutions for complex problems and translate requirements into a vision, architecture, and roadmap.
Must have:
  • 10+ years system software experience
  • Strong programming skills (C, C++, Go, Rust)
  • Distributed systems expertise
  • Excellent communication skills
  • Data science/deep learning knowledge
Good to have:
  • TensorFlow/PyTorch experience
  • Docker, Kubernetes, Slurm experience
  • CUDA/NCCL programming
  • HPC programming (MPI, OpenACC)
  • DGX Cloud, NVIDIA AI Enterprise experience
Perks:
  • Equity
  • Benefits

Job Details

NVIDIA is hiring engineers to scale up its AI Infrastructure. We expect you to have a strong programming background, a deep understanding of distributed systems, familiarity with software testing and deployment, and excellent communication and planning abilities. We also welcome out-of-the-box thinkers who can provide new ideas with strong at execution bias. Expect to be constantly challenged, improving, and evolving for the better. You and other engineers in this team will help advance NVIDIA's capacity to build and deploy leading infrastructure solutions for a broad range of AI-based applications that affect core data science. What are you waiting for if you're creative, passionate about what you do, and love having fun apply today!

We’re looking for a highly motivated, creative engineer with strong experience in system software to join the DGX Cloud Software Team. You will lead the architecture, design and implementation of our next generation DGX cloud clusters using latest technologies. On this team, you will do full stack deployment including hardware architecture, workload orchestration and application performance tuning. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement.

What you’ll be doing:

  • Lead technical activities for data centers with focus on hybrid deployments between cloud and on-prem

  • Providing expertise in infrastructure workflows, including hardware, workload orchestration and application tuning

  • Provide fast and creative solutions for complex problems and write effective, clear and reliable architecture specification

  • Translate requirements to vision, architecture and roadmap

  • Work with engineering teams across NVIDIA to ensure your software integrates seamlessly from the hardware all the way up to the AI training applications.

What we need to see:

  • Masters or PhD in Computer Science, Computer Engineering, Physics or equivalent experience

  • 10+ years of experience in this field.

  • Data Sciences, Deep Learning, or Machine Learning coursework

  • Ability to seamlessly shift between Linux system environments to Python programming

  • Programming skills in 1 or more high-level languages (C, C++,Go,Rust etc)

  • System-level experience with both hardware and software

  • Motivated self-starter with an equal balance of strong problem-solving skills and customer-facing communication skills

  • Strong design, coding, analytical, debugging and problem-solving skills

  • Passion for continuous learning and knowledge transfer. Ability to work concurrently with multiple groups locally and abroad in the organization

Ways to stand out from the crowd:

  • Experience with GPU deep learning and data sciences. Experience using TensorFlow, PyTorch or other DL framework. Experience working with Docker containers, Slurm, Terraform and Kubernetes

  • CUDA programming and NCCL experience. HPC programming experience including MPI, OpenACC, or other parallel programming tools

  • Hands-on experience with DGX Cloud, NVIDIA AI Enterprise AI Software, Base Command Manager, NEMO and NVIDIA Inference Microservices.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you are creative and autonomous, we want to hear from you!

The base salary range is 184,000 USD - 425,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

Blizzard Entertainment - Senior Data Scientist, Computer Graphics

Blizzard Entertainment

Irvine, California, United States (On-Site)
2 Months ago
NVIDIA - Senior Solutions Architect, Retail

NVIDIA

United States (Remote)
2 Weeks ago
NVIDIA - Senior CPU Implementation Methodology Engineer

NVIDIA

Santa Clara, California, United States (On-Site)
3 Weeks ago
Match Group - Sr. Software Engineer, Machine Learning Infrastructure

Match Group

Palo Alto, California, United States (Hybrid)
4 Months ago
Rackspace Technology - AWS Devops Engineer I - R-20532

Rackspace Technology

Gurugram, Haryana, India (Remote)
2 Months ago
Nagarro - Senior Engineer, Cloud

Nagarro

Bengaluru, Karnataka, India (On-Site)
4 Months ago
KBG Blockchain Game Studios - DevOps (Blockchain Gaming)

KBG Blockchain Game Studios

Thành Phố Hồ Chí Minh, Vietnam (On-Site)
7 Months ago
Hitachi - Azure Infra Consultant

Hitachi

Pune, Maharashtra, India (Remote)
4 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Intel Corporation - Senior IP Design Engineer (HBM Controller)

Intel Corporation

Center District, Israel (Hybrid)
2 Months ago
NVIDIA - Senior ASIC Physical Design Engineer - High Performance Designs

NVIDIA

Santa Clara, California, United States (On-Site)
1 Month ago
NVIDIA - Senior System Software Engineer - GPU Virtualization

NVIDIA

Pune, Maharashtra, India (On-Site)
1 Month ago
ByteDance - Research Scientist Graduate (Foundation Model, Video Generation) - 2025 Start (PhD)

ByteDance

San Jose, California, United States (On-Site)
3 Months ago
Luxoft - Regular Data Engineer

Luxoft

(Remote)
2 Months ago
Zazz - Machine Learning Engineer

Zazz

(Remote)
2 Days ago
Rackspace Technology - Presales Data Science Architect – AWS Cloud

Rackspace Technology

Mexico City, Mexico (On-Site)
3 Months ago
Meta - Software Engineer (Leadership) - Machine Learning

Meta

Burlingame, California, United States (Remote)
3 Months ago
Dream Sports - ML Engineer

Dream Sports

Mumbai, Maharashtra, India (On-Site)
2 Months ago
Impact Analytics - R&D Architect/Sr. Architect - Artificial Intelligence

Impact Analytics

Bengaluru, Karnataka, India (On-Site)
5 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Santa Clara, California, United States

Global Step - Chief Marketing Officer

Global Step

Texas, United States (Remote)
3 Months ago
Electronic Arts - Senior User Acquisition Specialist

Electronic Arts

Los Angeles, California, United States (On-Site)
2 Weeks ago
Discord - Senior Data Scientist, Analytics

Discord

San Francisco, California, United States (Remote)
1 Month ago
Trek - Assembler

Trek

Sacramento, California, United States (On-Site)
2 Weeks ago
ION - Application Support Engineer (Trading Systems)  - 5882

ION

New York, New York, United States (On-Site)
4 Months ago
ByteDance - Software Engineer, Architecture and Infrastructure

ByteDance

Seattle, Washington, United States (On-Site)
3 Months ago
Epic Games - Senior Environment Artist

Epic Games

Cary, North Carolina, United States (On-Site)
1 Month ago
Onward Search - Marketing Project Manager

Onward Search

Westwood, Massachusetts, United States (Hybrid)
3 Months ago
Lionsgate Games - Intern, FAST Channels

Lionsgate Games

Santa Monica, California, United States (On-Site)
1 Month ago
Samsung Semiconductor - Senior Staff Engineer, high speed analog

Samsung Semiconductor

San Jose, California, United States (Hybrid)
3 Months ago

Get notifed when new similar jobs are uploaded

DevOps Jobs

Microsoft - Senior Engineering Manager – CI/CD Engineering

Microsoft

Hyderabad, Telangana, India (On-Site)
4 Weeks ago
Intrepid Studios,  Inc  - DevOps Engineer (Kubernetes & Cloud Services)

Intrepid Studios, Inc

San Diego, California, United States (On-Site)
6 Months ago
ION - Cloud Engineer/Architect (DevOps)

ION

London, England, United Kingdom (On-Site)
4 Months ago
Scanline VFX - Senior DevOps Engineer

Scanline VFX

Seoul, South Korea (Hybrid)
1 Week ago
N-iX - Senior DevOps Engineer (with Java or Go Background)

N-iX

Netherlands (Hybrid)
5 Days ago
Northern Trust - Manager, Infra Info Svcs

Northern Trust

Pune, Maharashtra, India (On-Site)
3 Months ago
SmileGate - [CTO본부] DBA 담당

SmileGate

Seongnam-si, Gyeonggi-do, South Korea (On-Site)
1 Month ago
PlayStation Global - IT Systems Engineer-Cloud

PlayStation Global

Carlsbad, California, United States (On-Site)
3 Months ago
Info Stretch - Lead Data Engineer

Info Stretch

Chennai, Tamil Nadu, India (On-Site)
3 Months ago
Virtuos - Lead Software Engineer

Virtuos

Singapore (On-Site)
3 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.


Yokne'am Illit, North District, Israel (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (On-Site)

United States (Remote)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Bengaluru, Karnataka, India (Hybrid)

Bengaluru, Karnataka, India (Hybrid)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug