LLM client for Swift. Sophon is a Swift Package for iOS 18+, macOS 15+, and Mac Catalyst that talks to Google Gemini, OpenAI and every OpenAI-compatible endpoint (Groq, Mistral, OpenRouter, DeepSeek, Qwen, GLM, Kimi, Doubao), and Anthropic's Claude through one kernel: schema-constrained structured output written once and encoded in each provider's dialect, configurable retry policies, model catalogs that carry lifecycle and free-tier metadata and fall back automatically when a provider retires a model, live model listing, and lenient decoding for the JSON that LLMs actually return. The request/retry/decoding kernel was extracted from three production iOS apps that each carried it as copy-pasted code, and it ships in all three today (see Used in).
智子, the proton-sized intelligence from The Three-Body Problem: it observes and reports.
| Product | Depends on | Contents |
|---|---|---|
SophonCore |
nothing | LLMSchema (one schema, two dialects), LLMMessage / LLMPart, LLMRetryPolicy + LLMRetryLoop + LLMHTTP, LLMModelPreset / LLMModelInfo / LLMModelStore / LLMCatalogAudit, LLMProviderConfiguration (key + availability helpers), the LLMClient protocol, LLMErrorCopy (the error copy every provider shares), LLMDecoding (lenient LLM JSON decoding), LLMJSONExtractor (fence stripping, brace extraction, truncation repair), SophonKeychain, SophonLogger, UIKit-gated LLMImageEncoder |
SophonGemini |
SophonCore |
GeminiAPIClient, GeminiClientConfiguration, GeminiModel catalog, GeminiError, Gemini request/response DTOs |
SophonOpenAI |
SophonCore |
OpenAICompatibleClient (Responses API and Chat Completions), OpenAIEndpoint, OpenAIError, catalogs + endpoint presets for OpenAI, Groq, Mistral, OpenRouter, DeepSeek, Qwen, GLM, Kimi, Doubao |
SophonAnthropic |
SophonCore |
AnthropicAPIClient, AnthropicClientConfiguration, AnthropicModel catalog, AnthropicError, Messages API DTOs |
Platforms: iOS 18+, Mac Catalyst 18+, macOS 15+ (Foundation surface only; image encoding is #if canImport(UIKit)).
Add to your Package.swift:
dependencies: [
.package(url: "https://github.com/Luminoid/Sophon.git", from: "0.3.0"),
]Sophon is listed on the Swift Package Index, which hosts the DocC API reference for SophonCore, SophonGemini, SophonOpenAI, and SophonAnthropic.
Each app defines one configuration and one shared client per provider it ships. Only the Keychain account is required; Sophon supplies the model defaults (see Letting Sophon choose models).
import SophonGemini
extension GeminiClientConfiguration {
static let myApp = GeminiClientConfiguration(
keychainAccount: "com.myapp.geminiAPIKey",
logHandler: { level, message in MyLogger.log(level, message) }
)
}
extension GeminiAPIClient {
static let shared = GeminiAPIClient(configuration: .myApp)
}OpenAI-compatible providers share one client type; the catalog picks the endpoint, and providers with a China split take a region:
import SophonOpenAI
extension DeepSeekClientConfiguration {
static let myApp = DeepSeekClientConfiguration(keychainAccount: "com.myapp.deepSeekAPIKey")
}
extension QwenClientConfiguration {
static let myApp = QwenClientConfiguration(
keychainAccount: "com.myapp.qwenAPIKey",
endpoint: QwenModel.endpoint(region: .china)
)
}
let deepSeek = DeepSeekAPIClient(configuration: .myApp)
let qwen = QwenAPIClient(configuration: .myApp)Claude:
import SophonAnthropic
let claude = AnthropicAPIClient(configuration: AnthropicClientConfiguration(keychainAccount: "com.myapp.anthropicAPIKey"))One-call structured generation is the same call on every client. The schema is written once; each client encodes it in its provider's dialect (Gemini responseSchema, OpenAI strict json_schema or json_object, Anthropic output_config):
struct Extraction: Decodable { let title: String }
let result = try await GeminiAPIClient.shared.generateStructured(
Extraction.self,
label: "extract",
prompt: promptText,
schema: .object(properties: ["title": .string()], required: ["title"])
)Multi-turn plain text:
let reply = try await claude.generateText(
label: "followUp",
contents: [
LLMMessage(parts: [.text("You are a botanist.")], role: .system),
LLMMessage(parts: [.text("Is my monstera overwatered?")], role: .user),
]
)Image-heavy flows use the closure-based send so retries can re-encode smaller images and swap models:
let result = try await client.send(MyResult.self, label: "identify") { variant in
let parts = try await client.encodeImages(images, variant: variant)
return try client.buildRequest(parts: parts, promptText: prompt, apiKey: apiKey, modelID: variant.modelID, responseSchema: schema)
}Apps that let users pick a provider hold any LLMClient; it covers generateStructured, generateText, listModels(), and the current model ID.
Every catalog carries metadata (info: lifecycle with shutdown dates, free-tier membership, generation, pricing, image input, sampling support) and three adoption helpers, so an app can decide once how much to delegate:
// Explicit presets: the app owns the roster and updates it by hand.
availableModels: [.gemini36Flash, .gemini37Flash, .gemini38Flash]
// Every non-deprecated preset: a Sophon update adds new models and drops retired ones.
availableModels: GeminiModel.current
// Narrower rosters: by generation, free tier, or image support.
availableModels: GeminiModel.current(minimumGeneration: 3)
availableModels: GLMModel.currentFreeTier
availableModels: OpenAIModel.currentWithImageInput
// Omitted: Sophon's recommended default and fallback plus `current`.
defaultModel: .recommendedDefault
fallbackModel: .recommendedFallbackA stored selection outside the app's roster walks the catalog's successor chain (the provider's documented replacements) and only then falls back, so pruning presets never strands a user's stored choice. The roster always wins: a default or fallback outside availableModels resolves to the first offered preset, so the client never runs a model the app's own picker can't show. .custom(id) always passes through, and listModels() returns what the provider serves right now for pickers that want the live roster (listFreeModels() on OpenRouter). LLMCatalogAudit.violations(in:) checks a catalog's invariants; every Sophon catalog passes it in tests. Every catalog also exposes keyHintURL, where the user gets a key.
Recommended defaults today: Gemini 3.8 Flash (fallback 3.5 Flash-Lite), GPT-5.6 Luna (GPT-5.4 mini), Claude Sonnet 5 (Haiku 4.5), Groq gpt-oss-120b, Mistral Small 4, OpenRouter Nemotron 3.5 Lightning (free), DeepSeek V4.1 Flash, Qwen Flash, GLM-4.7 Flash, Kimi K2.6, Doubao Seed 2.1 Turbo. Where a provider has a permanent free tier, the recommended default is on it; OpenRouter's fallback is the paid openrouter/auto router on purpose, because its free roster rotates and the safety net must not.
Users bring their own API key, so the out-of-box experience matters. As of 2026-09-23 (each catalog's freeAccess carries the same facts):
| Provider | Free access | Notes |
|---|---|---|
| Gemini (pricing) | Permanent free tier | 3.x Flash and Flash-Lite models on an AI Studio key with no billing account; 3.1 Pro is paid-only, and the 2.5 models answer only projects that used them before 2026-09-18 |
| Groq (limits) | Permanent free plan | Every production model, for example gpt-oss-120b at 30 RPM, 1K RPD, 200K TPD; a payment method only to upgrade to the Developer tier |
| Mistral | Permanent free plan | Free mode (the default for new accounts): every model within limits shown in the console; data may be used for training unless you opt out |
| OpenRouter (limits) | Permanent free models | IDs ending in :free: 20 RPM, 50 RPD (1,000 RPD after a one-time $10 credit); the roster rotates |
| GLM (pricing) | Permanent free models | glm-4.7-flash and glm-4.6v-flash at a low, undocumented concurrency (glm-4.5-flash only on Z.ai) |
| Qwen (quota) | New-user quota | 1M tokens per model for 90 days, on the China site (Beijing) and the international site (Singapore) |
| Doubao | New-user quota | 500K tokens per model for new Volcengine Ark accounts |
| DeepSeek (pricing) | Signup credit | Off-peak hours bill at half price |
| Kimi | Signup credit | ¥15 voucher on the China platform after real-name verification, valid three months, not usable on kimi-k3 |
| OpenAI (pricing) | Trial credit | A rate-limited free trial tier for new accounts (no published amount); complimentary daily tokens for organizations that opt into data sharing, on GPT-5 and older models |
| Claude | Trial credit | A small free credit for new Console accounts; usage is prepaid afterwards |
Not in the catalogs but reachable through OpenAIEndpoint(...) plus .custom(id): SiliconFlow, Tencent Hunyuan, Baidu Qianfan, MiniMax, Xiaomi MiMo, NVIDIA NIM, and local servers.
Retry behavior is a parameter, not a baked-in default. Set it per app in the configuration, or override per call.
.default |
.minimal |
|
|---|---|---|
| Attempts | 3 | 3 |
| Backoff | 0.8s base, 6s cap, deterministic jitter | 1s base, 6s cap, no jitter |
Retry-After header (seconds or HTTP date) |
honored (clamped) | ignored |
| Retired model | retries the call on the fallback model | fails the call |
| Transport failure | re-encodes images smaller | no re-encode |
| Oversized request (HTTP 413) | re-encodes images smaller, once, and only when the request builder can (the generate* conveniences take pre-encoded parts, so they fail at once instead of re-sending the same body) |
fails the call |
Either way, a retired model persists a reset of the stored selection to fallbackModel, so the user's next call succeeds. Gemini treats every 404 as a retired model; the OpenAI-compatible and Claude clients reset only when the error body names the model, so a mistyped base URL never touches the user's selection.
Every call logs one .info line on success through the configuration's logHandler (provider, label, model, elapsed time, HTTP status, and the attempt count when it retried); retries and failures log where they happen, and model listings log how many models came back. Raw model output only ever appears at .debug, which the default handler logs privately.
GeminiError, OpenAIError, and AnthropicError descriptions resolve from localized strings (en, es, zh-Hans, zh-Hant) with app-neutral wording: provider-flavored lines live in each provider target, and the cases every provider shares (parse failure, rate limit, network error, image encoding, empty input, truncation, oversized request, locked Keychain) come from LLMErrorCopy in SophonCore. The OpenAI-compatible copy names the endpoint ("Groq API key not set"). A key the Keychain refuses to read (device locked) is apiKeyInaccessible, not "API key not set"; configuration.apiKeyStatus() tells the two apart. Apps that want feature-specific copy ("Gemini returned a trip we couldn't read") map the cases at their feature layer.
The Example/ directory contains a small iOS catalog app exercising the package end to end: a provider list (all fourteen endpoints, China regions included, grouped by how each can be used for free) with per-provider API key and model settings (the stored key shows its last four characters, can be revealed in full, and copies to a device-local clipboard that clears after a minute), catalog metadata, and live model listing (OpenRouter lists its free roster, and any listed ID can be picked as a custom model), a provider-and-model picker on the generation pages, schema-constrained structured output, multi-turn chat with a menu of test messages that each carry a catch (a letter count, a negative constraint, an instruction hidden in the data, a two-turn recall), and the offline JSON extractor (no API key needed). Key actions and Sophon's own lines land in one os.Logger subsystem (dev.luminoid.sophon.example), the way an app bridges logHandler to its logger. It uses XcodeGen to generate the Xcode project:
cd Example
xcodegen generate
open SophonExample.xcodeprojbrew bundle # install swiftlint + swiftformat + xcodegen
make setup-hooks # wire pre-commit lint + format
make check # SwiftLint --strict + SwiftFormat --lint
make build-strict # swift build with warnings as errors (library + tests)
make test # xcodebuild, iOS simulator (canonical)
make test-host # swift test (fast, Foundation-only surface)The same gates run on GitHub Actions for every push and pull request (.github/workflows/ci.yml).
Sophon carries the Gemini integration in three App Store apps for iPhone, iPad, and Mac:
| App | What it is |
|---|---|
| Plantfolio | Plant care: AI plant identification, seasonal watering schedules, collections (site) |
| Petfolio | Pet care: health logs, vet visits, medication schedules, Family Sharing (site) |
| TripDays | Collaborative travel planner: itineraries, paste-to-fill travel links, shared trips, expense splitting (site) |
MIT. © Luminoid. See LICENSE and CHANGELOG.
- Monolith: CLI that scaffolds iOS apps, Swift Packages, and Swift CLIs (Sophon was scaffolded with it)
- Everything else at luminoid.dev