Together AI is an AI infrastructure platform for running, fine-tuning, evaluating, and deploying open and proprietary models through APIs, serverless inference, dedicated compute, and developer tools.
Serverless inference across many text and media models
Dedicated endpoints, provisioned throughput, and GPU clusters
Fine-tuning, embeddings, reranking, code sessions, and storage
Explore with AI
About
Together AI is a developer platform for running open-source and commercial AI models through APIs, serverless inference, dedicated endpoints, GPU clusters, fine-tuning, code execution, storage, and production infrastructure. It gives teams access to many model families without operating the full serving stack themselves.
The product is useful for developers, AI startups, and engineering teams building applications on open models. Together AI supports text, image, audio, video, embeddings, reranking, dedicated inference, provisioned throughput, fine-tuning, GPU clusters, code interpreter sessions, and file-system storage.
Together AI pricing is usage-based rather than a simple subscription. Costs vary by model, input tokens, output tokens, media generation, GPU hour, storage, code sessions, fine-tuning, and reserved capacity. That makes it powerful for production builders, but cost monitoring is essential.
For FutureStack, Together AI belongs in Developer Tools because it is infrastructure for AI products. It is best for teams that need model APIs, open-model inference, dedicated endpoints, fine-tuning, GPU capacity, and scalable AI deployment options.
Use Cases
Model Inference
Run supported AI models through APIs and hosted infrastructure for applications and production workloads.
Model Fine-tuning
Customize supported open models for domain-specific behavior and application requirements.
AI Application Hosting
Use serverless or dedicated compute options to deploy model-powered applications and services.
Model Evaluation
Test and compare models and configurations before choosing an approach for production AI systems.
Key Features
Serverless inference across many text and media models
Dedicated endpoints, provisioned throughput, and GPU clusters
Fine-tuning, embeddings, reranking, code sessions, and storage
Usage-based pricing across tokens, GPUs, media, and infrastructure
We may earn commissions from links to support our work. Learn more.
Explore with AI
Pricing summary
Together AI uses usage-based pricing. Serverless models are priced per token or media output, GPU clusters are priced per GPU hour, and fine-tuning, storage, code sessions, and dedicated capacity have separate rates.
Pricing & Plans
Popular
Serverless Inference
Pay per model usage across text, image, audio, and video APIs.
Custom
Per-token pricing
Media model pricing
Embeddings and reranking
Pay as you go
Dedicated Inference
Reserve custom hardware for predictable serving performance.
Custom
GPU hourly rates
Guaranteed performance
Autoscaling options
Custom models
GPU / Fine-Tuning
Train, tune, and run models with GPU and fine-tuning infrastructure.
Custom
GPU clusters
Fine-tuning
Code sessions
Storage pricing
Other pricing notes
Rates vary by model, modality, GPU, storage, and deployment type.
Teams should estimate usage before choosing between serverless, dedicated, and reserved capacity.