Oopla
Spotlight can search. The terminal can act. Most people live stuck between those two. Oopla sits in that gap: a floating command bar that understands plain English and either launches something immediately or plans and runs a sequence of tools.
What it is
You know what you want done, but typing it as a shell command is not how you think, and Spotlight will not open Chrome, go to a job posting, and tailor a resume for you.
Hit Cmd+Shift+Space. A borderless SwiftUI window appears near the top of the screen, Spotlight-style. Type a command.
Two paths kick in. Instant local: apps, files, calculator, system settings via LocalSearchIndex. AI path: Claude returns a structured JSON ActionPlan, SafetyEvaluator checks it, then ToolRegistry executes step by step. CommandOrchestrator owns that routing.
A note on sources: the root README still describes AI as stubbed behind MockPlanner. The running app does not — OoplaApp wires ClaudePlanner and VisionPlanner with a live Anthropic key. "Token cap" here means API max_tokens (1024 text / 2048 vision) plus conversation windows (store 20 turns, send the last 10), not a separate budget system. A comment in HotkeyService still mentions Option+Space; the real hotkey is Cmd+Shift+Space.
Architecture
CommandBarView
Cmd+Shift+Space floating bar
A borderless SwiftUI window appears near the top of the screen, Spotlight-style. Status bar and hotkey share the same toggle. The window is accessory — no Dock icon — and dismisses on resign.
Jump to details →CommandBarViewModel
Input, results, and run state
Owns the typed query, local result list, and the handoff into the orchestrator. Instant local suggestions update as you type; Enter kicks off plan-and-run.
Jump to details →CommandOrchestrator
Local vs AI routing
shouldEscalateToClaude decides whether to short-circuit into a local open or call Claude. Vision queries, multi-step markers, mutating words, and weak local matches escalate. Conversation history (last 10 of 20 turns) rides along.
Jump to details →LocalSearchIndex · ClaudePlanner · VisionPlanner
Apps, files, plans, screenshots
Local search covers apps, Spotlight files, calculator, and settings. ClaudePlanner returns a JSON ActionPlan. VisionPlanner grabs a ScreenCaptureKit frame and sends it multimodal when the query needs eyes.
Jump to details →SafetyEvaluator
safe · confirmationRequired · destructive
Every tool declares a SafetyLevel. The evaluator walks the plan and keeps the strongest level. Safe plans run immediately; confirmation and destructive plans wait in pendingConfirmation.
Jump to details →ToolRegistry → ExecutionView
Step-by-step tool execution
Registered tools launch apps, open files, browse, compose mail, export PDFs, and more. ExecutionView streams pending → running → success/failure for each step.
Jump to details →OoplaApp builds the graph once: registry, planners, capture service, orchestrator. The window is accessory (no Dock icon), borderless, floating, dismiss-on-resign. Status bar and hotkey share the same toggle.
init() {
let registry = ToolRegistry.makeDefault()
let apiKey = EnvLoader.get("ANTHROPIC_API_KEY") ?? ""
let capture = ScreenCaptureService()
_orchestrator = StateObject(
wrappedValue: CommandOrchestrator(
searchService: LocalSearchIndex(),
planner: ClaudePlanner(apiKey: apiKey),
visionPlanner: VisionPlanner(apiKey: apiKey, captureService: capture),
registry: registry
)
)
}The confidence gate
Not every keystroke should hit an LLM. Local opens should feel instant. API calls cost money and add latency.
shouldEscalateToClaude(query:candidates:) in CommandOrchestrator makes the call. Vision queries always escalate. Multi-step markers (and, then, after), mutating words (create, delete, email), and web keywords also escalate. Everything else checks the top local candidate.
Apps only short-circuit when the normalized query (after stripping open / launch) is at least three characters and closely matches the app title. Files and folders need a score of 0.8 or higher. Otherwise Claude plans. buildLocalPlan only runs when the gate is sure.
if top.kind == .app {
guard normalized.count >= 3 else {
logger.log("Escalation reason: normalized query '\(normalized)' too short for local match")
return true
}
let title = top.title.lowercased()
// We do NOT accept normalized.hasPrefix(title) — a short app name
// like "screen" must not match "what is my screen showing".
let isMatch = title == normalized
|| title.hasPrefix(normalized)
|| title.contains(normalized)
if isMatch { return false }
return true
}
if top.kind == .file || top.kind == .folder {
if top.score >= 0.8 { return false }
return true
}The Raycast false open
That three-character rule exists because of a real bug. Typing a single letter like r still matched Raycast in local search. The gate treated it as a high-confidence open and launched the wrong app. The fix was simple: refuse local app matches under three characters, and stop accepting normalized.hasPrefix(title) so a short app name cannot swallow a long natural-language query. LocalSearchIndex.fuzzyMatch also refuses the reverse contains check for the same reason.
/// NOTE: We intentionally do NOT check `q.contains(t)` (whether the query
/// contains the app name), because a long natural-language query like
/// "what is my screen showing" would then match any short app name whose
/// letters happen to appear in the query string (e.g. "R.app" matches
/// because "screen" contains the letter 'r').
func fuzzyMatch(query: String, target: String) -> Bool {
let q = query.lowercased()
let t = target.lowercased()
if t.contains(q) { return true }
return t.components(separatedBy: .whitespaces).contains { $0.hasPrefix(q) }
}Giving it eyes
Some commands need the screen. "Explain what's on my screen." "Tailor my resume for this job."
requiresVision looks for phrases like that. When true, VisionPlanner takes over. ScreenCaptureService grabs a native-resolution frame with ScreenCaptureKit. The JPEG goes to Claude as a multimodal message alongside a tool-schema prompt. The model returns the same JSON plan shape as text-only mode.
let pixelWidth = Int(CGDisplayPixelsWide(display.displayID))
let pixelHeight = Int(CGDisplayPixelsHigh(display.displayID))
config.width = pixelWidth > 0 ? pixelWidth : display.width
config.height = pixelHeight > 0 ? pixelHeight : display.heightThe screenshot hallucination
This failed once in a boring way. Early encoding squashed screenshots hard: 1 MB cap, shrink by 25% each pass down toward 600px. Claude misread on-screen text and confidently invented names, schools, and URLs. The fix was two parts. Capture and encode at higher fidelity: CoreGraphics pixel dimensions, a 4 MB budget, and a long-edge target of 1568 instead of aggressive downscaling. And prompt rules that say only report text you can actually read; if it is blurry, say "I can't read that clearly" instead of guessing.
/// Max compressed JPEG size sent to the API (4 MB).
private static let maxImageBytes = 4_194_304
/// Claude vision reads best around this long-edge size.
private static let targetLongEdge: CGFloat = 1568
private func encodeImage(_ image: NSImage, maxBytes: Int = maxImageBytes) -> String? {
let longest = max(image.size.width, image.size.height)
let scale = longest > Self.targetLongEdge ? (Self.targetLongEdge / longest) : 1.0
// ... scale, then JPEG compress from 0.9 down toward 0.5 until under budget
}Accuracy rules for reading the screen:
- Only state text you can actually read clearly in the screenshot.
- NEVER guess or infer details like names, URLs, institutions, or numbers.
- If text is too small or blurry, explicitly say "I can't read that clearly"
rather than guessing.Small text can still fail. The model is just less willing to invent an answer now.
Conversation, not one-shot
Follow-ups matter. After a screen explain, "what about that diagram?" should not force a full restart.
CommandOrchestrator keeps a conversation array of ConversationTurns, capped at 20. Planners receive the last 10. Vision mode stores activeConversationScreenBase64 and reuses it unless the query asks to look again (look again, now, current screen). That avoids a fresh capture on every follow-up and saves tokens. Recapture is still available when the screen actually changed.
let history = Array(conversation.suffix(10))
if requiresVision(query: query), let vp = visionPlanner {
let shouldRecapture =
shouldRefreshScreenContext(query: query) || activeConversationScreenBase64 == nil
let (visionPlan, usedBase64) = try await vp.createPlanWithVisionContext(
for: query,
candidates: searchResults,
attachedFileContext: attachedContext,
history: history,
preferredScreenBase64: activeConversationScreenBase64,
shouldRecapture: shouldRecapture
)
activeConversationScreenBase64 = usedBase64 ?? activeConversationScreenBase64
plan = visionPlan
}API limits stay explicit: 1024 tokens for ClaudePlanner, 2048 for VisionPlanner. Enough for a plan and an explanation, not an essay. "Token cap" here means API max_tokens plus those conversation windows, not a separate token-budget system.
The safety layer
Every tool declares a SafetyLevel: safe, confirmationRequired, or destructive. SafetyEvaluator.evaluate walks the plan and keeps the strongest level. Safe plans run. Confirmation and destructive plans land in pendingConfirmation until the user approves.
func evaluate(plan: ActionPlan) -> SafetyDecision {
var strongest: SafetyLevel = plan.requiresConfirmation ? .confirmationRequired : .safe
var reason = "All actions are safe."
for step in plan.steps {
guard let tool = registry.tool(named: step.toolName) else { continue }
switch tool.safetyLevel {
case .destructive:
strongest = .destructive
reason = "\(tool.name) is destructive."
return SafetyDecision(level: strongest, reason: reason)
case .confirmationRequired where strongest != .destructive:
strongest = .confirmationRequired
reason = "\(tool.name) requires confirmation."
case .safe:
break
default:
break
}
}
return SafetyDecision(level: strongest, reason: reason)
}That covers folder creates, mail compose, and PDF export. Resume tailoring adds a content rule, not just a confirmation gate. Both planner prompts say: use only real experience from the user's resume. Never invent jobs, skills, or qualifications. pdf_create_tool still requires confirmation before writing Resume_[Company].pdf to disk.
Confirmation before side effects.
Creating folders, composing mail, and writing PDFs all declare confirmationRequired. The plan can look correct and still wait for an explicit approve, because a wrong path on disk is worse than one extra click.
Never invent the user's experience.
Resume tailoring is exactly the failure mode where a capable model produces a plausible career. The prompt constraint is the same idea as Bopple's "never fabricate the user's work" rule: correctness pressure alone does not stop next-token prediction from filling gaps. The instruction has to say inventing is not allowed.
What's still rough
Several tools are stubs. Move, rename, and zip return "Not yet implemented." Notes, calendar, email draft, and DND are mocked. Mail send opens mailto: and cannot attach files natively.
The hotkey is Cmd+Shift+Space on purpose, to avoid fighting Spotlight. You need an Anthropic API key and Screen Recording permission for vision. Input Monitoring helps the global hotkey fire while other apps are frontmost. Tiny UI text still challenges the vision path even after the resolution fix.
private func isHotkey(_ event: NSEvent) -> Bool {
let flags = event.modifierFlags.intersection(.deviceIndependentFlagsMask)
// Cmd+Shift+Space — no macOS system conflicts.
return flags == [.command, .shift] && event.keyCode == 49
}How I built it
I owned the architecture and product calls: local vs AI routing, vision only when needed, conversation reuse, safety before execution, resume rules that refuse fabrication. I used AI tools heavily for Swift and SwiftUI implementation. I can walk through every decision above from the code, including the Raycast false open and the screenshot hallucination, because those failures shaped the current gates.