
How do I set it up?
Ask Hercules to build you an app that uses AI! Hercules will use Hercules AI Gateway by default e.g.Setup Hercules AI Gateway
What can I build with AI Gateway?
- Chatbots: Customer support chatbot, AI coach
- Content generation: Write descriptions, emails, summaries
- Image generation: Create images from a text prompt
- Voice and audio: Turn text into speech, transcribe audio and video to text
- Analysis: Extract insights from text, categorize data
- Summarization: Condense long documents or articles
- Personalization: Generate custom recommendations
Can I generate images and audio?
Yes. The AI Gateway handles images and audio through the same setup, with no extra API keys.- Image generation: Create images from a text prompt
- Text to speech: Turn written text into spoken audio, with a choice of voices
- Speech to text: Transcribe audio and video files into text
Add image generation
Add text to speech
Add transcription
Can my app make decisions, like sorting or flagging content?
Yes. Use AI Decisions with Jev to classify, route, moderate, and score content. It is faster and cheaper than a chatbot model for these tasks.What models can I choose?
Any current model from Anthropic, Cloudflare Workers AI, Google AI Studio, or OpenAI. The gateway passes the model string straight through to the provider, so new releases from these providers work as soon as the provider ships them. There is no allowlist to wait on. The table below is a starting point, not a complete list. Extra large and large models are best for complex work. Small models are faster and cost less. Tell Hercules the model you want by name, or use the model string when working with the API directly (see the technical FAQ below).What model does Hercules use by default?
By default, Hercules AI Gateway uses GPT-6 Luna by OpenAI. It’s fast, cheap, and very accurate.How do I choose another model?
If you’d like to use another model, just tell Hercules the model and/or provider you want to use:Use Claude Haiku by AnthropicUse Gemini 3.8 Flash by Google
Can I use a model that is not in the table?
Yes. Use the provider’s own model id after the provider prefix, for exampleanthropic/claude-fable-5-1 or openai/gpt-6-astra. If the provider serves the model, the gateway serves it. If the provider rejects the id, the gateway returns the provider’s error.
Workers AI model ids keep their own @cf/ namespace, so the full string has two parts after the prefix: workers-ai/@cf/qwen/qwen3.8-27b. Any model in the Cloudflare Workers AI catalog works this way.
Models from providers other than Anthropic, Cloudflare Workers AI, Google AI Studio, and OpenAI are not available.
What input/output is supported?
How is billing handled?
AI Gateway usage is billed separately through Hercules Cloud. It’s a separate line item from the Hercules App Builder Agent (what you use to build your app). We bill the model provider’s rate plus a 7.5% platform fee. For instance, if you use $10 of OpenAI, you will be billed $10.75 of Hercules Cloud. Enterprise plans can have a custom platform fee. Contact sales@hercules.app for details.Can I see what each AI request costs?
Yes, for most models, when your app uses the Responses or Messages format. Their text responses include what the request cost, in dollars, including the 7.5% platform fee. It is the same amount billed to your Hercules Cloud balance. Ask Hercules to show it in your app:Show AI cost per answer
Do I need my own API keys?
No. This is all managed by Hercules out of the box through theHERCULES_API_KEY.
If you want to manage billing yourself, you need to ask Hercules to remove the Hercules AI Gateway, then configure the API key and model yourself.
Additional FAQ
What's the difference between AI Gateway and the Hercules App Builder Agent?
What's the difference between AI Gateway and the Hercules App Builder Agent?
The Hercules App Builder Agent builds your app. AI Gateway lets your users use AI features inside your published app.They are also currently billed separately. Hercules App Builder Agent bills against your monthly AI credits. Hercules AI Gateway bills against your Hercules Cloud credits.
How does the Hercules AI Gateway work technically?
How does the Hercules AI Gateway work technically?
Use the OpenAI SDK with Hercules configuration:The
HERCULES_API_KEY is automatically available in your app’s environment.Which API formats does it support?
Which API formats does it support?
- Chat Completions (OpenAI SDK, shown above): the default
- Responses API (OpenAI SDK):
openai.responses.create({ model: "openai/gpt-6-luna", input }) - Messages API (Anthropic SDK), for
anthropic/...models:
usage.cost: what the request cost in US dollars, including the 7.5% platform fee. When streaming, it is on the final chunk that carries usage.Chat Completions responses do not currently include usage.cost. Estimate their cost from the usage token counts at the model’s price plus the 7.5% platform fee, and treat your Hercules Cloud usage as the exact amount.How do I generate images, speech, and transcripts?
How do I generate images, speech, and transcripts?
Use the same client. Switch the method and model:
Is there rate limiting?
Is there rate limiting?
Yes. Hercules handles rate limiting automatically. If you need higher limits, please contact
support.
Can I stream responses?
Can I stream responses?
Yes. Hercules can set up your chatbot to stream responses in real-time as they come through
How do I monitor AI Gateway usage?
How do I monitor AI Gateway usage?
View usage in your billing dashboard under Hercules Cloud.
Is there a size limit on AI requests?
Is there a size limit on AI requests?
Yes. Each chat or text request can be up to 8 MB. Audio transcription and image edit uploads can be up to 50 MB.Files sent inside a request, like a PDF or an image, grow by about a third when they are encoded, so a single request can carry roughly 6 MB of file. Larger requests are rejected with “Request body is too large”.For a large PDF, split it into smaller batches of pages and send each batch as its own request. You can ask Hercules to set this up for you.
What's the difference between this and AI Image Generation?
What's the difference between this and AI Image Generation?
AI Gateway image generation runs inside your published app, for your users, at runtime. The AI Image Generation feature is Hercules creating images for your app while you build it (logos, banners, illustrations).