For iOS developers

AI in your iOS app, at one flat price

Per-token pricing means every message your users send adds to your bill, and a good week on the App Store can turn into a bad invoice. Our plans are a flat monthly price with no token meter. Your AI cost stays the same whether your users send a hundred messages or a hundred thousand. When they all arrive at once, requests wait their turn instead of costing more.

See plans Jump to the Swift sample

What you get

Getting through App Review

Apple has no special approval for AI providers. What it reviews is your app and how it handles your users' data and the AI's output. These are the parts of the App Review Guidelines that concern an app that talks to an AI service, and how each one is covered.

This is a summary to help you prepare, not legal advice, and Apple makes the final call on every review.

Never put your API key in the app

Anything in your app binary can be pulled out of it, and a leaked key lets anyone use your plan. Put a small proxy in between that holds the key. Here is a whole one as a Cloudflare Worker (the free tier is enough for most apps). It only forwards chat requests, caps answer length and keeps thinking off. Replace the token check with your own sign-in, or with Apple's App Attest, before you ship.

// A Cloudflare Worker between your iOS app and the API, so your API key stays on the server.
// Secrets (wrangler secret put): ASC_KEY = your sk-unl-... key, APP_TOKEN = a check of your own.
const API = "https://ascompute.com.au/v1/chat/completions";

export default {
  async fetch(request, env) {
    if (request.method !== "POST") return new Response("Not found", { status: 404 });
    // Replace this with your real check: your own sign-in token, or Apple App Attest.
    if (request.headers.get("X-App-Token") !== env.APP_TOKEN) return new Response("Unauthorized", { status: 401 });
    let body;
    try { body = await request.json(); } catch { return new Response("Bad JSON", { status: 400 }); }
    // Only what the app needs: its messages, streaming, and a cap on answer length.
    const upstream = await fetch(env.API_URL || API, {
      method: "POST",
      headers: { "Content-Type": "application/json", Authorization: `Bearer ${env.ASC_KEY}` },
      body: JSON.stringify({
        model: "qwen3.8-27b",
        messages: Array.isArray(body.messages) ? body.messages.slice(-20) : [],
        stream: body.stream === true,
        max_tokens: Math.min(Number(body.max_tokens) || 800, 2000),
        reasoning_effort: "none",
      }),
    });
    return new Response(upstream.body, {
      status: upstream.status,
      headers: { "Content-Type": upstream.headers.get("Content-Type") || "application/json" },
    });
  },
};

Swift: stream a reply

No package needed. This streams the reply into your UI as it is written, using URLSession and Swift concurrency (iOS 15 and later).

import Foundation

struct ChatMessage: Codable {
    let role: String
    let content: String
}

/// Streams a reply from your proxy (or any OpenAI-compatible chat endpoint), piece by piece.
func streamReply(to messages: [ChatMessage], endpoint: URL, headers: [String: String]) -> AsyncThrowingStream<String, Error> {
    AsyncThrowingStream { continuation in
        let task = Task {
            do {
                var request = URLRequest(url: endpoint)
                request.httpMethod = "POST"
                request.setValue("application/json", forHTTPHeaderField: "Content-Type")
                for (name, value) in headers { request.setValue(value, forHTTPHeaderField: name) }
                request.httpBody = try JSONSerialization.data(withJSONObject: [
                    "model": "qwen3.8-27b",
                    "stream": true,
                    "reasoning_effort": "none",
                    "messages": messages.map { ["role": $0.role, "content": $0.content] },
                ])
                let (bytes, response) = try await URLSession.shared.bytes(for: request)
                guard let http = response as? HTTPURLResponse, http.statusCode == 200 else {
                    throw URLError(.badServerResponse)
                }
                for try await line in bytes.lines {
                    guard line.hasPrefix("data: ") else { continue }
                    let payload = line.dropFirst(6)
                    if payload == "[DONE]" { break }
                    guard let json = try JSONSerialization.jsonObject(with: Data(payload.utf8)) as? [String: Any],
                          let choice = (json["choices"] as? [[String: Any]])?.first,
                          let delta = choice["delta"] as? [String: Any],
                          let text = delta["content"] as? String else { continue }
                    continuation.yield(text)
                }
                continuation.finish()
            } catch {
                continuation.finish(throwing: error)
            }
        }
        continuation.onTermination = { _ in task.cancel() }
    }
}
// In a SwiftUI view model: show the reply as it arrives.
let proxy = URL(string: "https://your-worker.your-account.workers.dev")!
var reply = ""
for try await piece in streamReply(to: [ChatMessage(role: "user", content: prompt)],
                                   endpoint: proxy, headers: ["X-App-Token": appToken]) {
    reply += piece
}

How many requests at once do you need?

A plan runs a set number of requests at the same time; more than that wait in a short queue. A short chat reply (about 250 tokens with thinking off) takes around 3 seconds at our typical speed, so one request slot can write roughly 1,000 replies an hour when busy. If your busiest hour carries about a sixth of the day's messages and each user sends 10 messages a day, one slot covers a few hundred daily active users. Long answers or thinking on use more time per reply. Start with one, watch the usage on your account page, and add more when replies start to queue.

Standard or Priority? Standard is the cheapest and uses spare capacity, so a reply can occasionally wait or be interrupted when the GPUs are busy. For an app in production, Priority reserves the capacity for you. You can start on Standard and upgrade from your account page at any time; the unused Standard time is credited.

Questions?

Email [email protected]. We read every message and are happy to help you get your app through review.

See plans