Prepared Tuesday, October 6, 2026 (PT). These are plans only. No app code was changed, no accounts were created, and no money was spent. Prices are in US dollars. "Per 1M" means per 1 million tokens (roughly 750,000 words). Anything I couldn't confirm from an official page is marked Unverified.
Bottom line
- Build order: (1) finish the Replit copy, (2) the Private switch, (3) Agents, (4) the Cloud computer. Number 4 is parked until you say "build it."
- Private switch (Plan 1): a second switch next to OpenRouter: Standard | Private 🔒. Private uses Tinfoil (strongest privacy you can check yourself) and Venice (only its private, hardware-enclave and uncensored models, never its "anonymized" Claude/GPT/Gemini). Every reply shows a badge: 🛡️ Hardware-verified (Tinfoil), 🔐 Hardware (Venice-verified), or 🤐 No-logs promise. Private mode never quietly falls back to OpenRouter.
- Cost: about $6–$11 a month if all ~1,200 messages went through Private, or about $2–$3 if a quarter of them do. Both services are pay-as-you-go with no subscription.
- Effort: about 3 working days of build time, plus about 20 minutes of your time to sign up.
- Agents (Plan 2): a new Agents tab where you create named helpers (for example Lead Writer, Estimator, Researcher, Marketing). Each one gets its own instructions, brain (model), Standard or Private, voice, and reference files. Type @Estimator in any chat and it answers in the thread with its own name and avatar. Agents handing work to each other comes in Phase 2.
- Cost: the same price per message as today, roughly $2–$6 a month for 300 agent replies.
- Effort: about 4–5 days for the first version, plus 2 days for files and 3–4 days for hand-offs.
- Cloud computer (Plan 3, parked): I recommend Kernel. It's pay-as-you-go with a free plan that includes $5 of use every month. You get a live view you can embed in the app (with a phone-sized 390×844 screen option), saved logins, and stealth and CAPTCHA handling at no extra charge. You don't pay while the browser sits idle. Runner-up: Browser Use Cloud, which is the cheapest per hour ($0.02) but charges $5/GB for proxies.
- Cost: about $3 a month for light use (20 tasks) and $30–$50 for heavy use (200 tasks), including the AI that drives the browser.
- Effort: about 5–6 days for the first version.
- You own all the code. All three plans are plain Flask code in your app. Every outside service sits behind a small swappable "adapter," and your data can be exported.
Contents
- Done so far
- Recommended build order
- Plan 1: the "Private" switch
- Plan 2: Agents
- Plan 3: Cloud computer with iPhone live view (parked)
- Sources
1. Done so far
- OpenRouter 8-model picker: tested. Gemini Flash (default), Flash Lite, Gemini Pro, Grok 4.7, DeepSeek V3.2, Claude Sonnet 5.5, Llama 4 Maverick and Mistral Medium. The server only accepts models on its own allowlist.
- Inworld TTS-2 voice: live on the private app. It falls back to Gemini "studio" voice, then to the phone's built-in voice.
- Security fixes done:
- App password in place.
- Hidden files are no longer served.
- Old password files wiped.
- Gemini key locked down.
- Business Gmail address removed from the health-check output.
- Replit copy: the
INWORLD_API_KEYandOPENROUTER_API_KEYsecrets are set. The new code is not applied yet; it's waiting for your test.
2. Recommended build order
| # | Step | Why this order | Rough effort |
|---|---|---|---|
| 1 | Finish Replit (apply the new code, you test it on your iPhone) | Everything else builds on a working copy, and the secrets are already in place | About half a day of build time, plus about 15 minutes of your testing |
| 2 | Private switch | Smallest change, clear value, and Agents can then offer "Private" as a brain option | About 3 days |
| 3 | Agents | Reuses the model picker, the Private switch and voices | About 4–5 days for the first version; about 9–11 days for all phases |
| 4 | Cloud computer (parked) | Biggest and riskiest piece; it needs the agent plumbing first | About 5–6 days for the first version |
"Days" means working days of build time by your builder (the AI that writes your code), including tests. Each plan below lists the few things you need to do yourself.
3. Plan 1: the "Private" switch
3.1 What you'll see (kept simple)
- One switch at the top of the chat:
Standard|Private 🔒. Standard is today's OpenRouter picker. Private swaps the model list for private models only. The switch remembers your choice on your phone. - The Private model list is short and plain-English. It's grouped and tagged so you never need to know model names:
- Everyday (default), Smart, Hardest, Uncensored, plus a "More private models" section that stays folded shut.
- Each line shows a badge: 🛡️ Hardware-verified, 🔐 Hardware (Venice-verified) or 🤐 No-logs promise, plus a price hint ($, $$, $$$).
- Every Private reply carries a small badge, for example "🔒 Tinfoil · Hardware-verified" or "🔒 Venice · No-logs promise", so you always know where an answer came from.
- A lock pill in the header ("🔒 Private") stays visible while Private is on, so you can't forget which mode you're in.
- Uncensored models get an "18+" tag and a one-time "Are you sure?" the first time you pick one.
- If Private isn't set up yet (no keys), the switch is greyed out and says "Not set up yet."
- Tap the badge to open a one-screen explainer of what each label means (see 3.2).
3.2 What the three labels mean (plain English)
| Badge | Provider | What it actually promises |
|---|---|---|
| 🛡️ Hardware-verified | Tinfoil | The model runs inside a sealed hardware "enclave." Before sending anything, our own server checks a cryptographic fingerprint of that enclave using Tinfoil's open-source SDK. If the check fails, nothing is sent. Tinfoil can't read your prompts (docs). |
| 🔐 Hardware (Venice-verified) | Venice e2ee-* models |
The model runs in a hardware enclave (run by Venice's partners NEAR AI and Phala). Venice's servers do the enclave check and tell us verified: true. We confirm that answer and a fresh random code before sending. In this basic "TEE" mode, Venice's relay still sees the text in transit (Venice TEE guide, privacy). |
| 🤐 No-logs promise | Venice "private" models | Venice and its hosts promise "zero data retention, contract-enforced." This is a policy promise, not hardware proof (privacy). |
Ranking for fallbacks (strongest first): Hardware-verified, then Hardware (Venice-verified), then No-logs promise.
3.3 Which models go in
Checked live on Oct 6, 2026, about 5:47 AM PT. Venice's public /models endpoint listed 128 text models: 70 "private" (12 of them TEE/E2EE) and 58 "anonymized." Tinfoil ids and prices come from its official catalog. Sources: Venice models API, Venice pricing docs, Tinfoil catalog.
Main list (what you see first):
| Slot | Provider · model | API id | Badge | Price in / out per 1M | Notes |
|---|---|---|---|---|---|
| Everyday (default) | Tinfoil · DeepSeek V4.1 Flash | deepseek-v4-1-flash |
🛡️ | $0.65 / $1.45 | Fast, reads images, 1M context. Marked "experimental"; weakest Tinfoil uptime (about 97.9%) |
| Smart | Tinfoil · GLM-5.3 | glm-5-3 |
🛡️ | $1.80 / $5.75 | Smartest Tinfoil model (AA index 44.8) |
| Hardest ($$$) | Tinfoil · Kimi K3 | kimi-k3 |
🛡️ | $4.00 / $20.00 | Hardest questions and images; slow |
| Smart (backup) | Venice · GLM 5.3 (TEE) | e2ee-glm-5-3-p |
🔐 | $1.75 / $5.50 | Same model as Tinfoil Smart, different host |
| Uncensored | Venice · Venice Uncensored 1.2 | venice-uncensored-1-2 |
🤐 | $0.20 / $0.90 | Venice tags it most_uncensored; reads images |
| Uncensored (hardware) | Venice · Gemma 4 26B Uncensored (TEE) | e2ee-gemma-4-26b-a4b-uncensored-p |
🔐 | $0.19 / $0.88 | The only uncensored model that runs in an enclave. Short memory (64K context) |
"More private models" (folded shut):
| Model | API id | Badge | Price in / out per 1M |
|---|---|---|---|
| Venice · Kimi K3 (TEE) | e2ee-kimi-k3-p |
🔐 | $3.75 / $18.75 |
| Venice · Qwen 3.8 27B (TEE, reads images) | e2ee-qwen3-8-27b |
🔐 | $0.47 / $3.53 |
| Venice · DeepSeek V4 Flash (TEE, cheapest) | e2ee-deepseek-v4-flash |
🔐 | $0.182 / $0.373 |
| Venice · Gemma 4 Uncensored (reads images) | gemma-4-uncensored |
🤐 | $0.1625 / $0.50 |
| Venice · Venice Role Play Uncensored | venice-uncensored-role-play |
🤐 | $0.50 / $2.00 |
| Venice · GLM 4.7 Flash Heretic (de-censored) | olafangensan-glm-4.7-flash-heretic |
🤐 | $0.07 / $0.40 |
| Optional: Venice · Grok 4.7 | grok-4-7 |
🤐 | $2.27 / $6.80 |
About Grok 4.7 on Venice: Venice labels it "private," but it almost certainly runs on xAI's own servers under a no-retention contract. Grok 4.7 is already in your Standard picker, so this is an open question for you.
Deliberately left out:
- Every Venice "anonymized" model: all Claude, GPT, Gemini, Qwen Max/Plus and Aion models. In that mode the original company still sees your prompt (Venice privacy).
- Two "uncensored" models that are actually anonymized:
qwen-3-6-plus("Qwen 3.6 Plus Uncensored") andabliteration-abliterated-model-large-v2("Abliterated Large V2"). Both are excluded. - Tinfoil's GLM-5.3 Flash: it's being removed on Oct 9, 2026 (changelog).
Keeping the list safe: the server keeps this list as an allowlist, so any id not on it is rejected. A small weekly check script compares the list against Venice's /models and Tinfoil's catalog. It warns if a model disappears, changes price, or (most important) if a Venice model's privacy flips to "anonymized."
3.4 How the server routes it
Both app copies (private box on :8787 and Replit) get the same new file, so they stay in sync. It's built from these parts:
- New module
private_providers.py. It holds the allowlist from 3.3. Each entry has: our own id (for exampletinfoil:glm-5-3,venice:e2ee-glm-5-3-p), provider, upstream model id, label, short name, badge, uncensored flag, reads-images flag, reasoning setting, and a price hint. /api/chance/modelskeeps its current output and addsprivate: {available, models}.availableis true only if the keys are present. Key values are never returned./api/chance/chataccepts"provider": "private"plus a private model id:- The server checks the id against the private allowlist. Anything else returns the existing "unknown model" error.
- It reuses what's already there: the rate limiter, the streaming slots, the time limit, the reply format the phone already understands (newline-delimited JSON), and refusal detection. The front end's reading code doesn't change.
- Tinfoil path:
- Uses the official
tinfoilPython SDK (v0.14.0, needs Python 3.10+) (PyPI, SDK docs). - One
TinfoilAIclient is created the first time it's needed, then reused. That's the moment the enclave check runs; if it fails, the reply is "Private check failed, nothing was sent." - Streams the reply.
- Sends
reasoning_effort="low"for GLM-5.3, Kimi K3 and DeepSeek. These models think at "max" by default, and those hidden thinking tokens are billed (reasoning guide). - Hidden thinking text is never sent to the phone.
- Not plain HTTPS to
https://inference.tinfoil.sh/v1. Tinfoil says that path skips verification (direct API).
- Uses the official
- Venice path:
- Plain OpenAI-style calls to
https://api.venice.ai/api/v1/chat/completions, with streaming. - Always sends
venice_parameters: {include_venice_system_prompt: false, strip_thinking_response: true, enable_web_search: "off"}. That stops Venice from adding its own personality prompt and keeps thinking text out of replies (Venice parameters). - For 🔐 models only: before first use, and again every 10 minutes, the server calls
GET /tee/attestation?model=…&nonce=<64 random hex>. It requiresverified: trueand a matching nonce, and refuses debug-mode enclaves. The result is cached. If the check fails, that model is refused (TEE guide).
- Plain OpenAI-style calls to
- Optional add-on (Phase 1b), Venice E2EE:
- Our server encrypts the prompt itself so Venice's relay sees only scrambled text. Only the enclave can read it.
- The method is documented (secp256k1 + AES-GCM, streaming required). It means about 6–10 more hours of build time and turns off web search and tools on those models (TEE guide).
- If built, those models could earn a stronger badge. Recommended later, not in the first version.
- Health check (
/api/chance/health): addsprivate: {tinfoil: "configured"/"missing", venice: "configured"/"missing"}. It never shows key values or emails. - Logs: prompt and reply text is never logged. That's already true today; the plan keeps it that way and adds a test.
3.5 Keys (secrets)
- Two new secrets:
TINFOIL_API_KEYandVENICE_API_KEY.- Replit: stored in Replit Secrets.
- Private box: stored in the server's environment, or in a key file under
~/.secrets/, the same way the Inworld key is handled.
- Keys never appear in code, chat messages, the web page, logs or the health output. Both are added to the existing "scrub secrets from any error text" list.
- You paste the keys yourself into Replit Secrets. For the private box, use the same secure path as the Inworld key, never a chat message.
3.6 When something goes wrong (fallback rules)
Golden rule: Private never silently switches to OpenRouter or to a weaker badge.
| Situation | What happens |
|---|---|
| Tinfoil model is down or slow | If you turned on "Try another model if one fails" (this setting already exists), the app tries the next model with the same badge: Everyday, then Smart, then Hardest. Otherwise you get buttons: "Try Smart" / "Try Venice (🔐)". |
| Tinfoil enclave check fails | Stop. Nothing is sent. Message: "Private check failed. Try again or pick another private model." |
| Venice enclave check fails | That model is refused. You're offered the Tinfoil model of the same strength. |
| Out of credit (HTTP 402/429 "insufficient quota") | "Private credit is used up: top up at Tinfoil/Venice." No fallback, because this is a money issue. |
| Bad key (401/403) | "The private key needs checking." Owner-only detail goes to the server log, with the key never shown. |
| A weaker badge is the only option | Shown as a button you tap, for example "Use Venice no-logs model instead?" It never happens automatically. |
3.7 Other privacy leaks to close (important)
- Voice:
- Speaking a Private reply with Inworld or Gemini voices sends the reply text to Inworld or Google.
- Plan: in Private mode the default voice is your iPhone's own built-in voice, which runs on the phone. A setting, "Use Inworld voice in Private mode," lets you opt back in, with a one-line warning.
- Chat history on the phone:
- History is stored in your phone's browser, as it is today.
- Plan: add an option, "Don't keep Private chats on this phone," which clears them when you close the chat.
- Learned notes and playbooks are still sent as instructions. They go to the private model, so that's fine, but they're included in the token count.
- Chance AI's own server still sees your text before it encrypts or forwards it. That's unavoidable, and why the app's own security fixes matter.
3.8 Tests
All automated tests use fake (mocked) network calls and fake keys. They never spend money or need real keys. They follow the style of the existing tests/test_openrouter.py.
- Allowlist:
- An unknown private id is rejected with a 400 error.
- An OpenRouter id sent with
provider: privateis rejected. - An anonymized Venice id (for example
claude-sonnet-5-5) is rejected.
- Models endpoint:
- Private shows as unavailable when keys are missing.
- Labels and badges are correct.
- Key values are never in the output.
- Tinfoil:
- The correct model and
reasoning_effort="low"are sent, with streaming on. - Thinking text is stripped.
- A failed enclave check means no request is made.
- The correct model and
- Venice:
include_venice_system_prompt: falseis sent.- Streaming is parsed correctly.
- 🔐 models:
verified:falseis refused, a nonce mismatch is refused, and the 10-minute cache expires.
- Fallback:
- Never falls back to OpenRouter.
- Same-badge fallback happens only when you've opted in.
- Weaker badges are offered, never automatic.
- Error messages: 401, 402, 429, 5xx errors and timeouts each show a friendly message.
- Secret scrubbing: a fake key is injected and the test checks it never appears in any response or log line.
- Model-drift script (run by hand or weekly): warns about missing ids, price changes, or privacy flipping to "anonymized."
- iPhone checklist (you, about 10 minutes):
- Flip the switch and check the header pill and the badges.
- Send one message per main-list model.
- Check that the voice defaults to the phone voice.
- Check that uncensored models ask for confirmation.
- One live smoke test with your real keys (about 1¢ total): one short message per model.
3.9 Monthly cost
Assumptions: 1,200 messages × about 6,000 tokens in and 400 out = 7.2M input + 0.48M output tokens a month. These are my calculations from the official prices in 3.3.
| If every message used… | Monthly cost |
|---|---|
| Tinfoil Everyday (DeepSeek V4.1 Flash) | $5.38 |
| Tinfoil Smart (GLM-5.3) | $15.72 |
| Tinfoil Hardest (Kimi K3) | $38.40 |
| Venice GLM 5.3 TEE | $15.24 |
| Venice Uncensored 1.2 | $1.87 |
| Venice Gemma 4 26B Uncensored TEE | $1.79 |
| For comparison: today's default, Gemini Flash on OpenRouter | $7.20 |
- Likely mix if everything went Private (60% Everyday, 15% Smart, 5% Hardest, 15% Uncensored, 5% Venice GLM TEE): about $8.55 a month.
- That's about $6.10 if half the input is served from cache (Tinfoil gives a cache discount).
- It's about $10.70 if each answer adds around 600 hidden "thinking" tokens.
- Range: about $6–$11.
- More realistic (a quarter of messages Private): about $2–$3 a month extra.
- Caps:
- Set a $20/month spend limit on the Tinfoil key (Tinfoil key docs).
- Set a daily USD limit (for example $2) on the Venice key (Venice key guide).
- Assumptions: the cache rate and thinking-token numbers are illustrative guesses, not measurements.
3.10 Optional third provider: Privatemode, now or later?
- What it is:
- Hardware enclave plus end-to-end encryption.
- Run by Edgeless Systems in Germany.
- Free tier: 5M tokens at sign-up plus 1M a month, no card needed, then pay-as-you-go in euros.
- Models: GLM-5.3 (€1.55 / €5.74 per 1M) and gpt-oss-120b, plus OCR and speech (pricing).
- The catch:
- For Python apps it needs its encryption proxy running as a separate Docker container next to the app. That's easy on the private box, but awkward on Replit.
- It mostly duplicates GLM-5.3, which Tinfoil and Venice already provide.
- Recommendation: later. Add it as a free backup for "Smart" on the private box once Plans 1–2 are stable, or sooner if Tinfoil's outages become annoying. About 4–6 hours of build time when you want it.
3.11 Sign-up steps you do yourself (about 20 minutes)
Nobody else creates accounts or enters cards for you.
Tinfoil:
- Go to tinfoil.sh, then sign up.
- Dashboard: Private Inference → Activate. Add a card through Stripe. You're charged only for usage, and you get $1 of free credit (get an API key).
- Copy the default API key.
- Set a spend limit on the key, for example $20/month. Also set the account-wide usage limit.
- Replit → Secrets: add
TINFOIL_API_KEY. Tell the builder it's there, without pasting the key into chat.
Venice:
- Go to venice.ai, then sign up.
- Open Settings → API (venice.ai/settings/api).
- Buy credits with a card ($10 is plenty to start). Credits never expire. No Pro subscription is needed for the API (pricing docs). Minimum card top-up: Unverified.
- Generate New API Key:
- Type: Inference Only.
- Set an Epoch Consumption Limit (USD per 24 hours, for example $2).
- Copy it once; it's shown only once.
- Replit → Secrets: add
VENICE_API_KEY.
3.12 Effort
| Piece | Build hours |
|---|---|
| Private allowlist, models endpoint, health output | 3–4 |
| Tinfoil route (SDK, streaming, reasoning, errors) | 4–6 |
| Venice route (streaming, parameters, TEE check plus cache) | 4–6 |
| Switch, badges, header pill, voice default, uncensored confirm (front end) | 5–7 |
| Fallback rules and error messages | 2–3 |
| Tests and model-drift script | 4–5 |
| Port to Replit, notes, your test session | 2–3 |
| Total | 24–34 hours (about 3–4 days) |
| Optional: Venice E2EE (Phase 1b) | +6–10 |
| Optional: Privatemode | +4–6 |
Your time: about 20 minutes to sign up and about 10 minutes for the iPhone checklist.
3.13 Open questions for you
- Is Tinfoil Everyday (DeepSeek V4.1 Flash) fine as the default Private model, or do you want Smart (GLM-5.3) as the default?
- Show the uncensored models at all? And to beta guests too, or only to you?
- In Private mode, should the voice default to the phone voice (most private), or keep Inworld?
- Should beta guests see the Private switch, or only you?
- Are the spending caps OK: Tinfoil $20/month, Venice $2/day?
- Include Venice Grok 4.7 ("no-logs promise," but it runs at xAI)?
- Add Venice E2EE now, or later?
- Automatically clear Private chats from the phone?
4. Plan 2: Agents
4.1 What an agent is
An agent is a named helper with a job. Each agent has:
- Name and avatar: an emoji plus a colour in the first version; photo upload comes later.
- Instructions: "what should this agent do?" (up to 8,000 characters).
- Brain: Standard (OpenRouter) or Private, plus a model from that list. Only allowlisted models can be chosen.
- Voice: an Inworld voice, a Gemini studio voice or the phone voice.
- Knowledge (optional): reference files such as price sheets, scripts or FAQs, and/or your existing playbooks.
- Reply length cap: to control cost.
4.2 Screens (wireframe, in words)
- Tab bar:
Chat·Agents·Settings. It sits at the bottom, within easy thumb reach on the iPhone 16 Plus. - Agents tab:
- At the top, a big "+ New agent" button, plus a row of template chips: Lead Writer, Estimator, Researcher, Marketing.
- Below that, one card per agent: round avatar, name, a one-line job description, a brain badge (for example "Gemini Flash" or "🔒 Smart"), and a price hint ($ / $$ / $$$).
- Tap a card to edit it.
- Edit agent screen (one scrolling page, big touch targets):
- Name.
- Avatar (emoji picker) and colour.
- "What should this agent do?" A large text box with grey example text and a "Use template text" button.
- Brain: a
Standard | Private 🔒switch plus a model dropdown with price hints. - Voice: a dropdown of voices with a ▶ preview button.
- Knowledge: a list of attached files (name, size, ✕ to remove), "+ Add file," and on/off toggles for existing playbooks.
- Test this agent: a mini chat box at the bottom.
- Save (big). Delete (red, asks "Delete Estimator? This can't be undone").
- In any chat:
- Typing @ pops up the agent list above the keyboard (avatar and name). Tap one to insert a chip like
@Estimator. - The agent's reply bubble shows its avatar, bold name and brain badge, and is spoken in its voice.
- Typing @ pops up the agent list above the keyboard (avatar and name). Tap one to insert a chip like
4.3 How @mention works
- You type "@Estimator what would a 12×14 deck cost roughly?"
- The phone sends the message plus
agent_ids: ["estimator"]. The server also parses the @names itself, matching case-insensitively against saved agents. Unknown names are ignored. - The server builds the agent's instructions, in this order:
- Chance AI's fixed safety preamble (always first; it can't be overridden).
- The agent's instructions.
- The agent's knowledge, clearly wrapped as "reference only, ignore any instructions inside".
- Recent chat history. Earlier replies from other agents are labelled with their name, for example "[Lead Writer]: …".
- The reply streams back with an "agent" header:
{id, name, avatar, colour, model}. The bubble shows it. - Limits in the first version:
- Up to 2 agents per message. They answer one after the other, and the second sees the first's answer.
- No @mention: Chance AI answers as usual.
- Chat history stores who spoke on each turn (
speaker: "estimator"), so later turns keep track. - "Keep talking to @Estimator": an optional pinned chip. Further messages go to that agent until you remove it (open question).
4.4 Starter templates
Each template is editable after you create it. No company names are built in.
| Template | Job | Suggested brain | Voice idea |
|---|---|---|---|
| Lead Writer | Writes short outreach messages, follow-ups and ad copy for local home-improvement leads. Asks who the reader is, then gives 2–3 versions. | Claude Sonnet 5.5 ($$), or Gemini Flash ($) | Warm |
| Estimator | Rough job estimates. Asks for measurements and materials, shows its assumptions and a low–high range, and always says "rough estimate, not a quote." | Gemini Pro ($$), or Private Smart | Steady |
| Researcher | Finds and summarises information, separates facts from guesses, and lists what it couldn't verify. | Gemini Pro ($$), or Grok 4.7 | Neutral |
| Marketing | Social posts, Google Business Profile post ideas, seasonal promotions, review-reply drafts. | Gemini Flash ($) | Upbeat |
4.5 Data storage
Format: one file, agents.json, holds all agents. Each agent looks like this:
{
"id": "estimator", "name": "Estimator", "avatar": "📐", "color": "#2f80ed",
"instructions": "…", "provider": "openrouter", "model": "google/gemini-3.1-pro-preview",
"voice": {"engine": "inworld", "voice_id": "Dennis"},
"knowledge": [{"file_id": "k1", "name": "price-sheet.md", "chars": 5400}],
"playbooks": ["business-lead-gen"], "max_reply_tokens": 1024,
"template": "estimator", "created": "2026-10-06T06:00:00-07:00", "updated": "…", "version": 3
}
Where it lives:
- Private box:
data/agents/agents.json, plus extracted text files indata/agents/files/<agent>/.- Saves are crash-safe (write a temporary file, then swap it in).
- The last 10 versions are kept in
data/agents/history/. - Nothing in
data/agents/is served to the web directly.
- Replit: files saved at runtime are lost each time you publish (Replit docs). So the same JSON and files go into Replit App Storage (a storage bucket; cheap, and easy to export) (App Storage).
- The Replit notes currently say "no database unless the owner asks," so this needs your OK (open question).
- One storage "adapter" with two back ends (
localandreplit), picked by a setting. The rest of the code doesn't care which one is used. - Export / Import button: downloads all agents and their files as one .zip. You always own your agents, and moving between hosts is just an upload.
Limits:
- Up to 25 agents.
- Up to 5 knowledge files per agent, 2 MB per upload. Accepted types: .txt, .md, .pdf, .docx (only the text is kept).
- Up to 60,000 characters of knowledge stored per agent.
- In the first version, up to 12,000 characters of knowledge are sent per reply. Phase 2 sends only the most relevant parts.
4.6 Server endpoints
All endpoints sit behind the existing password gate and use the same same-site checks as Settings.
| Method & path | What it does |
|---|---|
GET /api/chance/agents |
List agents (no file contents) |
POST /api/chance/agents |
Create (validates name, model allowlist, provider, voice, sizes) |
GET /api/chance/agents/<id> |
One agent |
PATCH /api/chance/agents/<id> |
Edit (bumps the version and keeps history) |
DELETE /api/chance/agents/<id> |
Delete (kept in history for undo by the builder) |
POST /api/chance/agents/<id>/files |
Upload a knowledge file (type and size checks, text extraction, warns if it looks like it contains passwords or keys) |
DELETE /api/chance/agents/<id>/files/<file_id> |
Remove a file |
GET /api/chance/agents/templates |
Starter templates |
POST /api/chance/agents/from-template/<template> |
Create from a template |
GET /api/chance/agents/export · POST /api/chance/agents/import |
Download or upload the .zip |
POST /api/chance/chat (extended) |
Accepts agent_ids. Stream events add agent_start / agent_end with name and avatar |
Phase 2: POST /api/chance/chat with hand-offs |
Stream adds handoff events ("Lead Writer → Estimator") |
4.7 Guardrails and cost controls
- Who can edit:
- Only the owner password (
BETA_PASSWORD) can create, edit or delete agents. - Guests (
BETA_PASSWORDS) can only @mention them. - This is new: the server will remember which kind of password signed in.
- Only the owner password (
- Brains are allowlisted: an agent can only use models from the Standard or Private allowlists.
- Private stays private: Private agents never fall back to OpenRouter. The Plan 1 rules apply.
- Safety preamble first: agent instructions can't remove Chance AI's base rules.
- Knowledge is "reference only":
- Files are wrapped so text inside them can't take over the agent (a defence against "prompt injection").
- On upload, a check (reused from the private app's "looks personal or secret" filter) warns about passwords, keys or personal data.
- Cost:
- Per-agent reply cap (default 1,024 tokens).
- Maximum 2 agents per message.
- Knowledge budget per reply.
- A monthly spend meter, built from the token counts OpenRouter, Tinfoil and Venice return, with warnings at 80% and 100% of a budget you set.
- The hard stop remains the credit limit on each provider key.
- Expensive brains show $$$ in the picker.
- Rate limits: each agent reply counts against the existing chat rate limit.
- Safe display: agent names and avatars are escaped before they're shown, so a name can't inject code into the page.
Cost example: 300 agent replies a month, each sending about 8,000 tokens (instructions plus knowledge plus history) and getting 500 back:
- Gemini Flash: about $2.36.
- Tinfoil GLM-5.3: about $5.18.
- Claude Sonnet 5.5: about $6.30.
4.8 Tests
- Create, edit, delete: each validates its input (name length, bad model ids, wrong provider, oversized instructions).
- Storage:
- Agents survive a restart.
- A crash mid-save doesn't corrupt the file.
- History keeps 10 versions.
- Export and import round-trip exactly.
- @mention parsing:
- Case-insensitive matching.
- Unknown names are ignored.
- At most 2 agents.
- Works with the chip and with plain typed text.
- Chat routing:
- The agent's instructions and model are used.
- The safety preamble comes first.
- Knowledge is wrapped and trimmed to the budget.
- Reply events carry name and avatar.
- Permissions: guest passwords can't edit; the owner can.
- Private agents never fall back to OpenRouter.
- Files:
- Wrong file type is rejected.
- Too-large files are rejected.
- Secret-looking content triggers a warning.
- Safe display: a name like
<script>shows as plain text. - Phase 2 (hand-offs):
- Hand-off limit (3 per message) enforced.
- No agent can hand to itself in a loop.
- "Stop" works.
4.9 Phases and effort
| Phase | What's included | Build effort |
|---|---|---|
| 1: first version | Agents tab, create/edit/delete, 4 templates, brain/provider/voice, @mention (1–2 agents) with name and avatar, local plus Replit storage, owner-only editing | 30–40 hours (4–5 days) |
| 1.5: knowledge | File upload and extraction, playbooks per agent, export/import .zip, spend meter | 12–16 hours (about 2 days) |
| 2: hand-offs | Agents can pass work along. An agent calls a handoff tool (models that support tools) or writes "HANDOFF: @Name: reason". The server runs the next agent, with at most 3 hand-offs per message, no ping-pong, a token budget per message, a visible "Lead Writer → Estimator" divider, and a Stop button. Also adds smarter knowledge search (sends only the relevant parts). |
20–30 hours (3–4 days) |
Dependency: "Private" as an agent brain needs Plan 1 first.
4.10 Open questions for you
- Who can create or edit agents: only you, or guests too?
- On Replit, is it OK to use Replit App Storage for agents (needed because files reset on every publish)?
- Want the "Keep talking to @Agent" pinned mode, or one @mention per message?
- Which 4 templates first? Are Lead Writer, Estimator, Researcher and Marketing right?
- Monthly agent budget for the spend meter (for example $10)?
- Avatars: emoji is fine for the first version, or do you want photo upload early?
5. Plan 3: Cloud computer the AI drives, with iPhone live view and "Take over" (parked)
Status: parked. Nothing will be built until you say "build the cloud computer."
5.1 What this means
The AI gets its own web browser running in the cloud:
- It can click, type and read web pages to do a task for you.
- You can watch it live inside Chance AI on your iPhone.
- You tap "Take over" to drive it yourself (for example to type a password or solve a puzzle), then hand it back.
- Logins you make are saved so the AI doesn't have to sign in every time.
5.2 Services compared (prices from their pricing pages, checked Oct 6, 2026)
| Service | Pay-as-you-go? | Real price | Live view you can embed | Human takeover | Saved logins | Stealth / CAPTCHA |
|---|---|---|---|---|---|---|
| Kernel ⭐ | Yes. Free "Developer" plan with $5 of use every month, then usage. No card needed to start | Visible browser $0.48/hour, only while active (idle is free); invisible "headless" browser $0.06/hour; no proxy charges (pricing) | Yes. An iframe link with a readOnly=true option; a phone-sized 390×844 screen is available (live view, viewports) |
Yes (interactive live view) | Profiles, plus "Managed Auth" (3 free connections) | Stealth and CAPTCHA solver included, even on the free plan |
| Browser Use Cloud | Yes, no subscriptions. $5 minimum top-up, credits never expire, $1 free | Browser $0.02/hour. Residential proxy $5/GB, on by default ($0.20/GB without proxy). Their own AI agent costs the model price plus 20% (pricing) | Yes (iframe; view-only possible) (live preview) | Yes (human-in-the-loop) | Profiles | Stealth and auto-CAPTCHA |
| Browserbase | Mostly subscription. Free plan = 1 browser hour total; Developer $20/month (100 hours, then $0.12/hour); Startup $99 | Proxies $12/GB after 1 GB on Developer (pricing) | Yes (interactive live view) | Yes | "Contexts" | CAPTCHA from Developer; "Basic" stealth |
| Steel.dev | Yes. "Launch" $0 + usage, $30 one-time credit (90 days); Scale $250/month | $0.10/hour, proxies $10/GB, CAPTCHA $3 per 1,000; a $10 deposit unlocks CAPTCHA and proxies (pricing) | Yes (embed live sessions) | Yes (human-in-the-loop guide) | Profiles | CAPTCHA yes; "Stealth Browser" is Enterprise-only. Sessions max 15 minutes on Launch |
| Hyperbrowser | Partly. Credits can be bought outright (they expire after 12 months) or come with plans | $0.10/hour, proxies $10/GB (pricing docs). Plans of $30 and $100 a month per third parties (Unverified) | Yes (liveUrl iframe, view-only option) (live view) |
Yes | Profiles | Stealth and CAPTCHA (paid tiers) |
| Anchor Browser | Not really. Free = $5 credit/month, but logged-in browsers and CAPTCHA need Starter at $50/month | $0.01 per browser + $0.05/hour + $8/GB proxy (pricing) | Yes (embedded live UI) | Yes | Profiles plus managed auth | CAPTCHA from Starter; full stealth on Growth ($2,000/month) |
| E2B Desktop | Yes. $100 one-time free credit, then per second; Pro $150/month for runs over 1 hour | About $0.17/hour for 2 CPUs / 4 GB (pricing) | A whole Linux desktop over VNC (works on iPhone, but clunky) (computer use) | Yes (it's a full desktop) | Pause/resume snapshots; no login-profile feature | None built in |
| Scrapybara | Was Free / $29 / $99 a month (third-party snapshot) | n/a | n/a | n/a | n/a | n/a. Site now says "Scrapybara built computers for agents. Now we're building Capy", and the pricing page returns 404. Avoid. |
5.3 If you meant a virtual phone
| Service | Price | Fit |
|---|---|---|
| Genymotion SaaS (virtual Android phone in a browser) | $0.06 per minute pay-as-you-go (= $3.60/hour), or $219/month per device for unlimited use (pricing) | Only worth it if a task needs an Android-only app. About 7× Kernel's hourly price |
| Limrun (cloud iOS and Android simulators) | No public price list. One comparison site quotes about $0.06/minute Android and $0.12/minute iOS (Unverified) (Limrun, comparison) | Built for app developers testing their own apps. iOS simulators can't install App Store apps, so it's not a "remote iPhone" |
Takeaway: for "do things on websites for me," a cloud browser is far cheaper and simpler. A virtual phone only makes sense for a specific Android app.
5.4 Recommendation: Kernel (runner-up: Browser Use Cloud)
Why Kernel fits you:
- No subscription. The free plan's monthly $5 credit covers light use completely, and after that you pay only for active minutes.
- Idle time is free. When the AI pauses for you, or you're deciding whether to take over, the browser costs nothing.
- Proxies, stealth and CAPTCHA are included, even on the free plan.
- Phone-sized screen (390×844), so the live view is readable on your iPhone 16 Plus.
readOnly=trueon the live view link gives a clean "Watch" vs "Take over" switch. The docs even note the Safari focus rule we need to follow.- Saved logins: profiles, plus Managed Auth that can re-login automatically.
- Ownership: browsers are controlled through the open standard Chrome protocol (CDP). The AI driver can be the open-source
browser-uselibrary (MIT license), running on your server with your OpenRouter models, and Kernel documents this exact pairing (Kernel + Browser Use). Kernel's browser images are open source too (Apache-2.0, GitHub). If Kernel ever disappoints, switching to Browser Use Cloud, Steel or Hyperbrowser means changing a single adapter.
When to choose Browser Use Cloud instead: if most tasks don't need residential proxies. Its browser hour is $0.02 against Kernel's $0.48, but proxies are $5/GB and on by default.
5.5 MVP design (one task, live view, Take over, saved logins)
What you'll see:
- A "🖥️ Computer" button in the chat. Type the task, for example "Find three local suppliers' prices for 2×6 pressure-treated boards and put them in a table."
- The live view opens in the chat (phone-sized). A status line underneath shows what the AI is doing, for example "Step 4 of ~20: typing in the search box." Two big buttons sit below it: Take over and Stop.
- When you tap Take over:
- The AI pauses at its next step.
- The view switches from watch-only to interactive.
- A "Type here" box appears below the view, which sends your typing straight into the cloud browser. This works around the iPhone keyboard sometimes not popping up inside embedded views.
- A "Give back to AI" button lets the AI continue, telling it "You took over; re-check the page."
- The AI asks before acting:
- Before anything that buys, pays, sends, submits, deletes or confirms, the AI stops and asks you to approve.
- A server-side check also blocks clicks on buttons with those words until you approve.
- Result: the answer appears in the chat as a normal reply, with an optional screenshot.
Saved logins ("Logins" screen):
- "Add a login": opens a fresh interactive browser at that site's login page, using a saved profile. You sign in yourself, including any 2-step codes. Tap Done and the profile is saved.
- Passwords never pass through Chance AI or the AI model.
- Later tasks load that profile and start already signed in.
- You can remove a login any time.
Server side:
- New secret:
KERNEL_API_KEY. - Endpoints:
POST /api/chance/computer/tasks(start)GET /api/chance/computer/tasks/<id>(status, steps, live link)POST …/takeover·POST …/resume·POST …/stop·POST …/typeGET/POST/DELETE /api/chance/computer/logins
- Task runner: each task runs in a background worker. A browser is created with the 390×844 screen, stealth on and the chosen profile. The
browser-useagent connects to it with an OpenRouter model (Gemini Flash by default) and reports each step back. - Live link security: the live-view link is a credential. It's served only to your signed-in session, never logged, and expires when the task ends.
- Limits:
- 1 task at a time.
- Maximum 15 minutes and 40 steps per task.
- Auto-stop when idle.
- Monthly budget meter.
- Hosting notes:
- Long tasks need an always-on server. On Replit that means a Reserved VM, not "Autoscale."
browser-useneeds Python 3.11+.
- Tests (mocked):
- Task start, stop and takeover state changes.
- Live link never in logs.
- Approval required for buy/send/submit actions.
- Step and time limits.
- Profile save and load.
- Plus an iPhone checklist for live view, Take over, typing and Give back.
5.6 Estimated monthly cost
Assumptions (illustrative, not measured):
- Each task is about 25 AI steps, each step sending about 6,000 tokens and getting 300 back.
- About 50 MB of web traffic per task.
- Light use = 20 tasks × 10 minutes. Heavy use = 200 tasks × 15 minutes.
| Light use (about 3.3 browser hours) | Heavy use (about 50 browser hours) | |
|---|---|---|
| Kernel browser time | $1.60, covered by the free $5 credit, so $0 | $24 − $5 credit = $19 |
| AI driver: Gemini Flash (about $0.14/task) | about $2.80 | about $28 |
| AI driver: Flash Lite (about $0.05/task) | about $1 | about $10 |
| Kernel total | about $1–$3 a month | about $29–$47 a month |
| Browser Use Cloud, browser plus proxies | about $5 (+AI) | about $51 (+AI); about $3 if proxies are off |
| Browserbase Developer | $20 (+AI) | $20 + proxies after 1 GB (+AI) |
| Genymotion virtual phone | about $12 | about $180 |
If you choose Claude Sonnet as the driver (smarter, about $0.38/task), add about $7.50 for light use or $75 for heavy use.
5.7 Effort and open questions
Effort: about 36–46 build hours (5–6 days) for the first version:
- Task runner and Kernel adapter: 10–12 hours
- Live view, Take over and Type box on iPhone: 8–10 hours
- Saved logins: 6–8 hours
- Approval guard and limits: 5–6 hours
- Tests and iPhone checklist: 7–10 hours
Your time: create a Kernel account (no card needed for the free plan) and add KERNEL_API_KEY, plus a 15-minute test.
Open questions for you:
- Did you mean a cloud browser (recommended) or a virtual phone for a specific app?
- What's the first task you'd want it to do?
- Which sites would you save logins for? (Some sites forbid automation in their terms, so we'd check each one.)
- Which AI should drive it: Gemini Flash (cheap) or Claude Sonnet (smarter, about 3× the cost)?
- Should it always ask before submitting anything? (Recommended: yes.)
- Monthly budget cap for this feature?
6. Sources (all accessed Oct 6, 2026)
Earlier research (this folder): research/tinfoil-deep-dive/index.md, research/tinfoil-vs-venice-verify-2026-10-06.md and research/privacy-llm-providers-2026-10-06.md.
App code (read only, not changed): the private copy's server.py, and the Replit copy's server.py and replit.md. These were used for the model list, the existing chat streaming, playbooks, and the "no database unless the owner asks" rule.
Tinfoil:
- Model catalog feed
- Python SDK · PyPI tinfoil 0.14.0
- Direct API (unverified path)
- Verification
- Reasoning
- API key and spend limits
- Changelog
Venice:
- Models API (live)
- Pricing docs
- Privacy modes
- TEE & E2EE guide
- API key and daily limits
- Docs index (venice_parameters)
Privatemode: pricing.
Replit: deployments (filesystem not persistent) · App Storage.
OpenRouter prices: saved model list research/or_models.json (Gemini 3.8 Flash $0.75/$3.75, Flash Lite $0.25/$1.50, Claude Sonnet 5.5 $2/$10, Gemini 3.1 Pro $2/$12).
Cloud browsers and computers:
- Kernel: pricing · pricing page · live view · viewports · profiles · Browser Use integration · kernel-images (Apache-2.0)
- Browser Use: pricing · live preview · human in the loop · profiles · browser-use library (MIT)
- Browserbase: pricing · docs index
- Steel: pricing and limits · docs index
- Hyperbrowser: pricing · live view · plan amounts (third party)
- Anchor: pricing · pricing page
- E2B: pricing · billing · computer use
- Scrapybara: home page · old plans (third party)
- Genymotion: pricing
- Limrun: home · comparison (third party)