For iOS developers
AI in your iOS app, at one flat price
Per-token pricing means every message your users send adds to your bill, and a good week on the App Store can turn into a bad invoice. Our plans are a flat monthly price with no token meter. Your AI cost stays the same whether your users send a hundred messages or a hundred thousand. When they all arrive at once, requests wait their turn instead of costing more.
See plans Jump to the Swift sample
What you get
- An OpenAI-compatible API. Chat completions with streaming, tool calls and JSON output, on Qwen3.8-27B. A Swift package that lets you set the host works, and so does plain
URLSession(below). - No per-token bill. Pay for how many requests your app runs at once, not for tokens. See the plans, or work out your app's numbers with the pricing calculator.
- Your users' messages are not kept. We do not log, store or read prompts or answers, and never train on them. Servers are in Australia. That is what Apple asks your AI provider to promise (see below), and it is in our privacy policy, which you can point to.
- Thinking off for snappy chat. Send
"reasoning_effort": "none"for quick replies, or leave thinking on for harder questions.
Getting through App Review
Apple has no special approval for AI providers. What it reviews is your app and how it handles your users' data and the AI's output. These are the parts of the App Review Guidelines that concern an app that talks to an AI service, and how each one is covered.
- Ask before sending personal data to an AI (5.1.2(i)). Apple requires you to say when personal data goes to "third parties, including with third-party AI" and to get explicit permission first. Show a short consent screen before the first message, such as: “Your messages are sent to Australian Standard Compute, an AI service in Australia, to write replies. They are not stored or used for training.”
- Your provider protects data as you do (5.1.1(i)). Your privacy policy must confirm that anyone you share user data with gives it the same protection. Name us as your AI provider and link our privacy policy: no logging of prompts or answers, no training, no selling, kept in Australia.
- App Privacy details. In App Store Connect, declare the data your app sends to the AI (usually User Content). We do not use it to track your users or for advertising.
- Filtering, reporting and blocking for chat (1.2, 4.7.1). Apps with chatbots need a way to filter objectionable output, report content and get a timely response. Set a system prompt that keeps the assistant on topic and refuses unsafe requests, keep a report button on AI messages, and act on reports. That part lives in your app.
- Age rating (2.3.6). Apple asks you to count AI assistant and chatbot features when you answer the age rating questions. Rate for what the assistant can be led to say, not just what it usually says.
- Charging for AI features (3.1.1). If users pay to unlock AI in your app, that goes through in-app purchase. A flat AI cost makes your margin predictable after Apple's commission.
- Network security. The API uses HTTPS with modern TLS, so App Transport Security needs no exceptions. Standard HTTPS is the kind of encryption the export compliance questions usually treat as exempt, but answer them for your own app.
- If users sign in, let them delete their account in the app (5.1.1(v)). That is about your app's accounts, not ours, but it is a common reason for rejection.
- China. Generative AI apps need a local licence on the China App Store. Most small developers leave China out of their availability.
This is a summary to help you prepare, not legal advice, and Apple makes the final call on every review.
Never put your API key in the app
Anything in your app binary can be pulled out of it, and a leaked key lets anyone use your plan. Put a small proxy in between that holds the key. Here is a whole one as a Cloudflare Worker (the free tier is enough for most apps). It only forwards chat requests, caps answer length and keeps thinking off. Replace the token check with your own sign-in, or with Apple's App Attest, before you ship.
// A Cloudflare Worker between your iOS app and the API, so your API key stays on the server.
// Secrets (wrangler secret put): ASC_KEY = your sk-unl-... key, APP_TOKEN = a check of your own.
const API = "https://ascompute.com.au/v1/chat/completions";
export default {
async fetch(request, env) {
if (request.method !== "POST") return new Response("Not found", { status: 404 });
// Replace this with your real check: your own sign-in token, or Apple App Attest.
if (request.headers.get("X-App-Token") !== env.APP_TOKEN) return new Response("Unauthorized", { status: 401 });
let body;
try { body = await request.json(); } catch { return new Response("Bad JSON", { status: 400 }); }
// Only what the app needs: its messages, streaming, and a cap on answer length.
const upstream = await fetch(env.API_URL || API, {
method: "POST",
headers: { "Content-Type": "application/json", Authorization: `Bearer ${env.ASC_KEY}` },
body: JSON.stringify({
model: "qwen3.8-27b",
messages: Array.isArray(body.messages) ? body.messages.slice(-20) : [],
stream: body.stream === true,
max_tokens: Math.min(Number(body.max_tokens) || 800, 2000),
reasoning_effort: "none",
}),
});
return new Response(upstream.body, {
status: upstream.status,
headers: { "Content-Type": upstream.headers.get("Content-Type") || "application/json" },
});
},
};
Swift: stream a reply
No package needed. This streams the reply into your UI as it is written, using URLSession and Swift concurrency (iOS 15 and later).
import Foundation
struct ChatMessage: Codable {
let role: String
let content: String
}
/// Streams a reply from your proxy (or any OpenAI-compatible chat endpoint), piece by piece.
func streamReply(to messages: [ChatMessage], endpoint: URL, headers: [String: String]) -> AsyncThrowingStream<String, Error> {
AsyncThrowingStream { continuation in
let task = Task {
do {
var request = URLRequest(url: endpoint)
request.httpMethod = "POST"
request.setValue("application/json", forHTTPHeaderField: "Content-Type")
for (name, value) in headers { request.setValue(value, forHTTPHeaderField: name) }
request.httpBody = try JSONSerialization.data(withJSONObject: [
"model": "qwen3.8-27b",
"stream": true,
"reasoning_effort": "none",
"messages": messages.map { ["role": $0.role, "content": $0.content] },
])
let (bytes, response) = try await URLSession.shared.bytes(for: request)
guard let http = response as? HTTPURLResponse, http.statusCode == 200 else {
throw URLError(.badServerResponse)
}
for try await line in bytes.lines {
guard line.hasPrefix("data: ") else { continue }
let payload = line.dropFirst(6)
if payload == "[DONE]" { break }
guard let json = try JSONSerialization.jsonObject(with: Data(payload.utf8)) as? [String: Any],
let choice = (json["choices"] as? [[String: Any]])?.first,
let delta = choice["delta"] as? [String: Any],
let text = delta["content"] as? String else { continue }
continuation.yield(text)
}
continuation.finish()
} catch {
continuation.finish(throwing: error)
}
}
continuation.onTermination = { _ in task.cancel() }
}
}
// In a SwiftUI view model: show the reply as it arrives.
let proxy = URL(string: "https://your-worker.your-account.workers.dev")!
var reply = ""
for try await piece in streamReply(to: [ChatMessage(role: "user", content: prompt)],
endpoint: proxy, headers: ["X-App-Token": appToken]) {
reply += piece
}
How many requests at once do you need?
A plan runs a set number of requests at the same time; more than that wait in a short queue. A short chat reply (about 250 tokens with thinking off) takes around 3 seconds at our typical speed, so one request slot can write roughly 1,000 replies an hour when busy. If your busiest hour carries about a sixth of the day's messages and each user sends 10 messages a day, one slot covers a few hundred daily active users. Long answers or thinking on use more time per reply. Start with one, watch the usage on your account page, and add more when replies start to queue.
Standard or Priority? Standard is the cheapest and uses spare capacity, so a reply can occasionally wait or be interrupted when the GPUs are busy. For an app in production, Priority reserves the capacity for you. You can start on Standard and upgrade from your account page at any time; the unused Standard time is credited.
Questions?
Email [email protected]. We read every message and are happy to help you get your app through review.