Skip to content
AI 도구

Modular 검토

Modular는 다양한 하드웨어 및 머신러닝 프레임워크에서 AI를 접근 가능하고 효율적으로 만들도록 설계된 통합 AI 엔진 및 플랫폼입니다.

shipped 2026년 7월 30일aifreemium
ai
Modular — product screenshot

핵심 포인트

1생성형 및 에이전트형 AI 소프트웨어 강화를 위해 2026년 7월 29일 Qualcomm에 인수되었습니다.
2Modular 26.4 릴리스(2026년 6월 18일)에서 최첨단 MoE(mixture-of-experts) 서빙을 도입했습니다.
3세 번째 자금 조달 라운드(2025년 9월 24일)에서 2억 5천만 달러를 유치하여 회사 가치를 16억 달러로 평가받았습니다.
4동일한 소프트웨어 플랫폼으로 Nvidia의 Blackwell B200 및 AMD의 MI355X GPU에서 최고의 성능을 달성했습니다.

Modular 소개

비즈니스 모델
Usage-Based (Pay Per Use)
사용량 기반 요금
Pay-per-token per token
본사
San Francisco, USA
팀 규모
100-250
플랫폼
Web, API, Cloud, On-premises
대상 사용자
AI developers, data scientists, enterprises focusing on AI deployment

요금제

Shared Endpoints
Pay-per-token / per-token
  • High-performance inference
  • No infrastructure management
  • Test easily
Dedicated Endpoints
Pay-per-minute / per-minute
  • Reserved NVIDIA and AMD GPUs
  • Flexible pricing
Custom Models
Pay-per-minute / per-minute
  • Deploy custom or fine-tuned models
  • Optimized infrastructure

비용 예시

  • Cost varies based on usage
API DocsGitHubOpen Source

사양

API 제공 여부

예, 공개 API

overview

Modular란 무엇인가요?

Modular는 AI 개발자, 애플리케이션 개발자 및 기업이 통합되고 고성능이며 이식 가능한 AI 추론 플랫폼을 구축할 수 있도록 지원하는 AI 네이티브 개발자 플랫폼 회사입니다. 하드웨어 복잡성을 추상화하여 개발자가 다양한 하드웨어에 AI 모델을 효율적으로 배포하고 확장할 수 있도록 하는 것을 목표로 합니다. Modular는 고성능, 이식 가능한 컴퓨팅을 위해 설계된 통합 AI 추론 플랫폼을 제공하며, GPU 커널에서 API 엔드포인트에 이르는 AI 파이프라인을 최적화합니다. 핵심 임무는 개발자가 코드 변경 없이 업계 최고의 GPU 및 CPU 성능으로 인기 있는 오픈 AI 모델을 실행할 수 있도록 하여 하드웨어 복잡성을 추상화하는 것입니다. 이 플랫폼은 OpenAI 호환 REST API를 통해 Hugging Face의 500개 이상의 AI 모델을 지원하며, 수천 개의 GPU 노드에 걸쳐 대규모 GenAI 추론 서비스를 확장하는 데 도움이 됩니다.

features

Modular의 주요 기능

Modular는 다양한 하드웨어 및 프레임워크에서 AI 개발 및 배포 효율성을 향상시키기 위해 설계된 포괄적인 기능 세트를 제공합니다.

  • 일관된 성능을 위한 통합 AI 추론 스택.
  • NVIDIA, AMD, Intel, ARM, Apple Silicon을 포함한 여러 하드웨어 플랫폼 지원.
  • Hugging Face의 30개 이상의 최첨단 모델과의 호환성.
  • 유연한 리소스 할당을 위한 종량제 요금 옵션.
  • 클라우드 환경 또는 온프레미스 인프라에서 유연한 배포 옵션.
  • 모델 서빙을 위한 OpenAI 호환 REST API.
  • 사용자 지정 작업 및 GPU 커널 개발을 위한 Mojo 프로그래밍 언어.
  • 몇 초 만에 더 빠른 개발 및 컴파일을 위한 Slim 툴체인.
  • GenAI 배포에서 처리량을 극대화하고 지연 시간을 최소화하기 위한 지능형 워크로드 라우팅.

use cases

Modular는 누가 사용해야 하나요?

Modular는 주로 다양한 하드웨어 환경에서 AI 배포를 최적화하고 확장하려는 AI 개발자, 애플리케이션 개발자 및 기업을 위해 설계되었습니다.

  • AI 개발자: Mojo로 사용자 지정 ops 및 GPU 커널 작성을 포함하여 다양한 하드웨어 및 클라우드 환경에서 고성능 AI 추론 및 배포를 위해.
  • 애플리케이션 개발자: PyTorch 및 TensorFlow 워크로드의 통합 AI 개발 및 배포를 위해 툴체인을 단순화하고 하드웨어 이식성을 달성합니다.
  • 기업: 대규모 생성형 AI(GenAI) 추론 서비스를 확장하고, GenAI 혁신을 가속화하며, AI 모델에 대한 공급업체 독립성을 달성하기 위해.

how to use

Modular 사용 방법

Modular를 사용하려면 개발자는 통합 플랫폼을 활용하여 AI 모델을 배포하고 최적화할 수 있습니다. 이 플랫폼은 단일 Docker 컨테이너 배포를 지원하며 모델 서빙을 위한 OpenAI 호환 API를 제공합니다.

  • 1웹 인터페이스 또는 API를 통해 Modular 플랫폼에 액세스합니다.
  • 2Hugging Face의 500개 이상의 사전 최적화된 AI 모델 중에서 선택합니다.
  • 3효율적인 설정을 위해 단일 Docker 컨테이너(1GB 미만)를 사용하여 모델을 배포합니다.
  • 4모델 서빙 및 애플리케이션 통합을 위해 OpenAI 호환 REST API를 활용합니다.
  • 5Mojo 프로그래밍 언어를 사용하여 사용자 지정 모델을 개발하고 성능을 최적화합니다.
  • 6다양한 하드웨어 백엔드에서 GenAI 추론 서비스를 모니터링하고 확장합니다.

pricing

Modular 가격 및 요금제

Modular는 유연성을 위해 설계된 사용량 기반 요금제를 갖춘 프리미엄 모델로 운영됩니다. 초기 탐색을 위한 무료 티어를 포함하며, 다양한 배포 요구 사항에 대해 별도의 가격을 제공합니다.

  • Shared Endpoints: 일반적인 사용을 위한 토큰당 지불.
  • Dedicated Endpoints: 일관된 고성능 액세스를 위한 분당 지불.
  • Custom Models: 특수 모델 배포를 위한 분당 지불.

Pros

  • +Mojo programming language combines Python's ease of use with C++/CUDA performance, addressing the 'two-language problem' in AI.
  • +Modular Platform (MAX) provides a unified AI inference stack, abstracting hardware complexity for portable, high-performance model deployment.
  • +Achieves significant performance gains, with Mojo benchmarks showing up to 35,000 times faster execution than Python in specific scenarios.
  • +Offers broad hardware compatibility, supporting NVIDIA, AMD, Intel, ARM, and Apple Silicon.
  • +Simplified community license allows free non-production commercial use and production use on CPUs and NVIDIA GPUs.
  • +Acquisition by Qualcomm on June 24, 2026, indicates strong industry validation and potential for future integration into edge AI.

Cons

  • Mojo is still in early development (Mojo 1.0 Beta 1 as of May 2026), leading to potential instability or evolving features.
  • Initial concerns existed regarding its closed-source nature and restrictive licensing, though the license has since been simplified.
  • Requires adoption of a new programming language (Mojo), which may present a learning curve for developers accustomed to other languages.
  • Some developers view Mojo as a 'toy for now' due to its early stage, suggesting more established languages like C++ or Rust for critical low-level programming.
  • The platform's full capabilities and long-term roadmap post-Qualcomm acquisition are subject to future announcements.

정책

가격 페이지

가격 보기

유사한 도구

Modular vs 경쟁사

Modular는 진정한 하드웨어 이식성 및 성능 최적화를 제공하여 기존의 하드웨어별 AI 소프트웨어 스택, 특히 NVIDIA의 CUDA 생태계에 도전하는 파괴적인 세력으로 자리매김하고 있습니다.

1

It's a deep learning compiler stack that optimizes models for various hardware backends, from CPUs to specialized accelerators, enabling efficient execution.

While Modular aims for a unified platform with a new language (Mojo) for performance, TVM focuses on compiling existing models for optimal performance on diverse hardware. You might need more integration work with TVM compared to Modular's more integrated platform approach.

2

It's a cross-platform inference engine that allows models from various frameworks (converted to ONNX format) to run efficiently on different hardware.

Modular aims to provide a high-performance engine and platform, potentially requiring adoption of Mojo. ONNX Runtime provides a standard format and runtime for existing models, offering broad compatibility and efficiency gains without changing your core ML framework, though you might need to convert your models to ONNX format.

3
NVIDIA Triton Inference Server

It's an open-source inference serving software that simplifies the deployment of AI models from various frameworks, offering features like dynamic batching, concurrent model execution, and GPU utilization.

Modular focuses on optimizing the AI engine itself for performance and development. Triton Inference Server focuses on optimizing the *serving* of those models in production, handling aspects like throughput and latency for deployed models. While Triton helps with efficient deployment, it doesn't offer the same level of low-level model optimization or a new development language like Mojo.

4
OpenVINO Toolkit

It's a comprehensive toolkit from Intel for optimizing and deploying deep learning models on Intel hardware (CPUs, GPUs, VPUs, FPGAs) and also supports other architectures.

Modular aims for hardware-agnostic efficiency with its platform and Mojo language. OpenVINO provides a robust set of tools specifically for optimizing and deploying models, particularly strong on Intel hardware, but requires more manual integration and doesn't offer a unified development language like Mojo.