Blog
How on-device AI actually works, what happens to the things you type into a cloud assistant, and how local models compare when you measure them honestly.
Start here
Each guide below is a complete introduction to one topic, with deeper articles beneath it.
All articles
Is Local AI Good Enough Yet? An Honest 2026 Assessment
Where on-device models have genuinely caught up, where the gap to frontier models is still large, and how to tell which side of the line your own work falls on.
Which iPhones Can Run Local AI? A Device Compatibility Guide
Which iPhone and iPad models can run on-device language models, how much RAM each model actually requires, and how to work out what your specific device will run.
Private AI for Writers: Drafts That Stay Unpublished
Unfinished work belongs in one place. How novelists, journalists, and academics use on-device models without their manuscripts leaving the device.
Offline AI for Travel: Translation and Guidance With No Signal
A local model on your phone works on the plane, in the tunnel, and in the country where you have no data. What it does well while travelling, and what it can't replace.
Asking AI the Questions You'd Never Type Into a Search Bar
Health, money, legal, and personal questions are where AI is most useful and most exposing. What makes these queries different, and how to ask them properly.
Private AI Journaling: Thinking Out Loud Without an Audience
Journaling with a model that responds is genuinely useful — but only if the space is actually private. How it works, how to set it up, and where the limits are.
What People Actually Use Private AI For
The real workflows where on-device AI beats a cloud assistant — not because the model is better, but because you can use it without editing yourself first.
How to Read an AI Privacy Policy: What the Terms Actually Permit
A field guide to the standard clauses in AI privacy policies — what 'we don't sell your data' leaves open, how training opt-outs really work, and the five questions that matter.
AI Apps That Don't Collect Your Data: How to Verify the Claim
Every AI app says it respects your privacy. Here are seven tests you can run yourself — in about ten minutes — that separate architectural guarantees from marketing language.
AI Chat With No Account, No Sign-Up, and No Subscription
Why almost every AI app wants an email address, what an account actually enables, and how on-device inference makes the whole requirement unnecessary.
ChatGPT vs a Local LLM: An Honest Everyday Comparison
Where a frontier cloud model genuinely beats a local one, where the gap has closed, and how to decide which to reach for — without pretending either side wins everything.
Private ChatGPT Alternatives: A 2026 Guide to AI That Can't Store Your Chats
Every category of private AI alternative, honestly assessed — local apps, self-hosted models, privacy-focused proxies, and enterprise modes — with the trade-offs each one actually makes.
On-Device Vision: How Your iPhone Reads Images Without Uploading Them
Multimodal models can describe, read, and analyse photos entirely on-device. Here's how vision models work, what they can and can't do at small sizes, and why the metadata question matters.
Thinking Mode: How Reasoning Models Work on Device
Reasoning models spend extra tokens working through a problem before answering. Here's what that buys you, what it costs, and when to turn it off.
Tokens Per Second: What AI Speed Numbers Actually Mean
Why benchmark speeds never match what you see, the difference between time-to-first-token and throughput, and the threshold where a local model stops feeling slow.
Context Windows Explained: How Much Your Local AI Remembers
What a context window is, why conversations slow down as they grow, what happens when you hit the limit, and how local models handle memory differently from cloud assistants.
LLM Quantization Explained: Why a 9B Model Fits in 5.6GB
Quantization is the compression that makes on-device AI possible. Here's what 4-bit weights actually mean, how much quality it costs, and why it makes models faster as well as smaller.
How LLMs Actually Run on Your iPhone
The complete technical explanation of on-device language models — unified memory, the Neural Engine, quantized weights, the KV cache, and why a 5.6GB model fits in your pocket at all.
The AI Privacy Crisis of 2026: What's Changed and What You Can Do
AI data practices are under more scrutiny than ever. Here's a clear-eyed look at what's actually happening in 2026, why your AI conversations are at risk, and the one architectural shift that changes the equation entirely.
On-Device AI: The Complete Guide to Running LLMs on Your iPhone
Everything you need to know about running large language models directly on your iPhone — how it works, which models run well, what the trade-offs are, and why it matters for your privacy.
What Is Apple MLX? The Framework Powering On-Device AI
Apple MLX is an open-source machine learning framework built specifically for Apple Silicon. Learn how its unified memory architecture enables fast, private on-device AI inference on iPhone and Mac.
Why Your AI Conversations Are More Sensitive Than You Think
People share things with AI that they wouldn't tell their doctor, lawyer, or closest friend. That's not a bug — it's a feature of the medium. But it creates privacy risks most users never consider.
The AI Privacy Checklist: 7 Things to Look For in Any AI App
Before you use an AI app for anything personal, run through these seven questions. They'll tell you more about your real privacy than any privacy badge or promise ever will.
AI Privacy: Why Your Conversations With ChatGPT Aren't As Private As You Think
Every message you send to ChatGPT, Gemini, or Claude travels to a corporate server, gets stored, and may train future models. Here's what actually happens to your data — and why on-device AI is the only architectural guarantee of privacy.
Best Local LLM Models for iPhone in 2026
A hands-on comparison of the 11 local LLM models you can run on iPhone — Qwen 3.5, Qwen 3, Llama 3.2, Phi-4 Mini, Ministral 3, and DeepSeek R1. Real sizes, real RAM requirements, real speeds.
What Happens to Your ChatGPT Conversations? A Data Privacy Deep Dive
ChatGPT stores your conversations by default, uses them to improve its models, and shares data with third parties. Here's exactly what happens to everything you type — and what you can do about it.
On-Device AI vs Cloud AI: Privacy, Speed, and Cost Compared
A direct comparison of on-device AI and cloud AI across privacy, latency, cost, and offline capability — so you can make an informed choice about where your conversations actually go.
How to Use AI Without Internet Access
A practical guide to using AI assistants with no internet connection — which apps work offline, how to set them up, and why going offline is often the smarter choice for sensitive conversations.
How to Run AI Completely Offline: A Practical Guide
A comprehensive guide to running AI locally on your devices — how offline inference works, what hardware you need, which use cases benefit most, and how to get started today without giving up your data.
Best Offline AI Apps for iPhone in 2026
A practical comparison of the best apps for running AI locally on your iPhone — what each one does well, which models they support, and how to choose the right one for your use case.
Open Source AI Models: Why They Matter and How to Use Them
Open source AI models like Llama, Qwen, Mistral, and Gemma give anyone the ability to run capable, auditable language models without paying API fees or trusting a third party with their data. This guide covers what open source means for AI, who builds the models, how licensing works, and how to actually run them on your own hardware.
Qwen 3 vs Llama 3: Which Runs Better on iPhone?
Qwen 3.5 4B and Llama 3.2 3B are the two most capable on-device language models for iPhone. Here's a direct comparison of their sizes, performance, thinking modes, and which tasks each handles best — with a clear recommendation for most users.
Small Language Models: Why Smaller Can Be Smarter
Small language models — models under 10B parameters — have gone from compromise to genuine alternative in two years. This post explains why they're improving faster than large models, how quantization and distillation work, and what it means for running capable AI privately on your phone.
Apple Intelligence vs Open Source On-Device AI: An Honest Comparison
Apple Intelligence and open source on-device AI like Cloaked both run AI locally, but they take fundamentally different approaches to models, privacy, and hardware requirements. Here's a fair look at the trade-offs.