It’s been a week since I wrote my first post on testing Apple’s on-device AI with Expo. In round one, I used OCR to pull data out of images for a few things I like to track in daily life. It was a good test. I was able to put together a simple iOS app with Expo quickly, run it on my phone, and use OCR to read my weight scale, MyFitnessPal, and my workout spreadsheet.

Today I’m on iOS 27, which lets apps send images straight to the on-device model. The goal is to see whether that’s a viable way to pull data out of images and log it accurately in the app.

First things first: I had some uncommitted changes for a bug with the OCR image labels I hit in round one. So I put up a draft PR to Expo, expo/expo#51375, to get that out of the way and on record. The app already had an image input mode from round one. It just never turned on. So next, I made a fresh build and tested on my phone.

Round Two Was Way Easier

No builds to set up and no Expo project to create. I kicked off a new build, scanned the QR code Expo gave me, and it was running on my phone.

Then the app’s model card said Vision: unsupported. My phone’s model won’t take images, so the app was stuck in OCR mode, same as round one. I reran the tests on iOS 27 anyway.

  • The workout screenshot came out as nicely structured data. The model clearly knew it was looking at a workout. But the numbers were wrong. It read the columns top to bottom instead of row by row, so Dips came out as 15/12/6 at 14 lb instead of 15/15/15 at 234 lb, and Chin ups got the next four numbers.
  • The scale photo (228.8 lb) just errored out. It called expo_ocr four times and hit ERR_TOOL_CALL_LIMIT, the cap on how many tool calls one request can make. At least it failed loudly instead of making up 250 like it did in round one.
The workout sheet screenshot above the first extracted rows, with Dips at 15, 12, 6 reps Extracted workout rows where the reps and weights follow the sheet's columns top to bottom The scale photo at 228.80 lb with an ERR_TOOL_CALL_LIMIT error after four expo_ocr calls

The good news: the image-label bug didn’t come back. Every call used the right label.

My Phone’s Model Can’t See

Digging into this, it turns out my phone’s copy of Apple’s model doesn’t support image input, and iOS reports that back to the app. Apple doesn’t ship the same model to every phone. It picks one that fits the hardware.

The model is called AFM, for Apple Foundation Model. Reading an image needs an extra component, an image encoder, which turns pixels into something the language model can understand. That makes the model bigger and hungrier for memory.

iOS 27 has two tiers: AFM 3 Core and AFM 3 Core Advanced. Advanced is the larger model and only runs on phones with 12 GB of RAM: the iPhone 17 Pro, 17 Pro Max, and iPhone Air, plus the new iPhone 18 Pro, 18 Pro Max, and iPhone Duo. I have an iPhone 16 Pro. What a bummer.

So my phone gets the Core model. That alone doesn’t explain it, though. More on that below.

To decide whether vision is available, expo-ai asks the model this:

SystemLanguageModel.default.capabilities.contains(.vision)

I added the device and iOS version to the app’s model card, plus a line saying why vision is off. It came back with “iPhone 16 Pro · iOS 27.0.1” and “Vision unavailable: this device’s system model does not report image input.”

The app's model card: iPhone 16 Pro, iOS 27.0.1, Vision unsupported, with the reason line in red

This felt like a dead end. It’s so lame. Am I supposed to go buy a new phone?

What’s Weird About It

I first read that Core is text-only, but that claim came from a GitHub issue, and it doesn’t hold up. Apple’s own post scores AFM 3 Core on image understanding, and a developer on an M2 Mac with 24 GB of RAM reports .vision as supported on the same base model. My iPhone 16 Pro, running that same base model, says no. I couldn’t find a single report of an 8 GB iPhone returning vision support, and Apple’s WWDC sessions and docs never say which devices get image input.

My best guesses, in order:

  1. A memory cutoff. Image input on the base model may need more than 8 GB of RAM. The M2 that worked has 24 GB, so this fits, but it’s an inference.
  2. An iPhone-specific restriction that Apple hasn’t mentioned.

I bet most people don’t know their phone has an on-device language model that any app can run, if you know how to get to it. I didn’t. As Apple keeps expanding it, I think there’ll be more and more use cases for it.

So Let’s Pivot

Instead of more extraction testing, I want to look at the models that actually ship on recent iOS devices. This is the kind of information Apple doesn’t put in its top-level marketing. It’s usually a snazzy ad about how the phone looks cool, the camera’s better, and the battery lasts longer. But I want to know about the models. That’s where a lot of the real advancement is, and we barely talk about it.

How the On-Device Model Has Changed

Apple doesn’t publish a per-device list of model capabilities, but you can piece the history together from its research posts, technical reports, and coverage.

Year / OSOn-device modelWhat changedImages through the developer API
2024, iOS 18AFM on-device, ~3B parametersFirst Apple Intelligence model, on iPhone 15 Pro and all iPhone 16s (8 GB RAM)No developer API yet
2025, iOS 26~3B, compressed to 2 bits per weightMemory savings across the board; 4,096-token contextNo. Text in, text out; images only reached the model through OCR/barcode tools
2026, iOS 27AFM 3 Core: 3B dense, improvedPreferred over the 2025 model on 45.6% vs 23.3% of text prompts, and on more than 61% of image-understanding comparisonsYes, on some devices
2026, iOS 27AFM 3 Core Advanced: 20B sparse, 1–4B active per requestFull model stored in flash, parts loaded into memory per request. Apple calls it “natively multimodal” but only names voices and dictation as usesNeeds 12 GB RAM: iPhone 18 Pro / Pro Max, iPhone Duo, iPhone 17 Pro / Pro Max, iPhone Air, M3 and newer Macs with 12 GB or more

That jump from 3B to 20B parameters is huge, even though only 1–4B of them run at a time.

A few more things I found:

  • The 2025 model could already see images. Apple’s 2025 technical report describes a ~300M-parameter image encoder. Apple used it for Visual Intelligence, like making a calendar event from a flyer, but kept it out of the developer API. So in round one, the model on my phone had an image encoder, and my app couldn’t use it.
  • The hardware split comes down to memory. The iPhone 16, 16 Pro, and 17 have 8 GB of RAM. The iPhone 17 Pro and Air have 12 GB. The A19 Pro still has a 16-core Neural Engine like the A18 Pro, and its AI gains come mainly from accelerators added to the GPU.
  • The iPhone 18 Pro is out, and it’s still 12 GB. The iPhone 18 Pro and Pro Max went on sale September 18, along with a foldable called the iPhone Duo. Their A20 Pro has a dual 16-core Neural Engine and, per Apple, 50% more memory bandwidth than the A19 Pro. RAM stays at 12 GB, which Apple doesn’t advertise but shows up in Xcode. Apple’s Apple Intelligence page lists the same advanced voice and dictation features for the 17 Pro and 18 Pro, so the new chip doesn’t add a model tier. The base iPhone 18 isn’t out yet. Reports put it in spring 2027.
  • Storage: Apple’s support note says Apple Intelligence on iOS 27 takes up to 14 GB of storage on the iPhone 17 Pro, 17 Pro Max, and Air, and up to 8 GB on the others. Coverage puts the 18 Pro in the 14 GB tier too. Apple lists devices, not RAM, but those are the 12 GB phones. That’s pretty substantial.
  • Context window: 4,096 tokens on base devices, which is the only number Apple documents. The developer who reported vision on an M2 says it’s 8,192 on M3 and newer Macs with 12 GB or more, and Apple’s WWDC sample prints 8192 as an example, so I’d expect the same on 12 GB phones.

How Hard Is It to Learn What Model Your Phone Has?

Pretty hard. Apple announces that the on-device model gains vision without saying which devices get it. The most reliable spec sheet is the runtime check, SystemLanguageModel.default.capabilities, which is exactly how my app was able to show me that notice. Apple’s reference docs for SystemLanguageModel don’t list that property yet, as far as I can find, even though it’s in the iOS 27 SDK. And Apple doesn’t surface any of this in Settings or a built-in app.

How AFM 3 Core Advanced Works

I wanted to learn more about the latest model, because the way it runs on a phone is clever.

The phone treats Core Advanced as a big library and only checks out the books a request needs. All 20 billion parameters sit in the phone’s flash storage. When a request comes in, a small router picks the 1–4 billion parameters the task needs, loads them into RAM once, and then the phone runs what is effectively a small, dense model.

Claude helped me build the film below in Remotion. It follows one nutrition screenshot through the phone, from the screen down to the chips and back out. The layout is stylized, and the memory and speed numbers are estimates.

My RTX 4080 Super doesn’t need that trick for a 20B model, because the whole thing fits in its 16 GB of VRAM (video RAM, the memory on the graphics card).

The router comes from an Apple paper called Instruction-Following Pruning (IFP). It picks the slice once per request instead of once per token, because pulling weights out of flash for every token would be way too slow.

Why It Needs 12 GB

The active slice is small, around 0.25–2 GB by my rough math. What needs the 12 GB is everything around it: iOS, your open apps, the model’s working memory for your conversation (the KV cache), and room to load experts without pushing apps out. I’d guess the 3B Core model stays resident too, but Apple doesn’t say. On an 8 GB phone, that margin isn’t there.

Compared With My RTX 4080 Super

My PC has an RTX 4080 Super with 16 GB of VRAM and 64 GB of system RAM. Here’s how it stacks up:

iPhone 18 ProMy PC
Fast memory12 GB, shared by CPU, GPU, and Neural Engine16 GB GDDR6X VRAM on the GPU
Fast memory bandwidth~115 GB/s (my estimate)~736 GB/s (about 6× more)
Slow tierFlash storage, a few GB/s (my guess)64 GB system RAM over PCIe 4.0 x16, ~32 GB/s
Fast-to-slow bandwidth ratioroughly 30–60×roughly 23×
Powera few wattsthe 4080 Super alone draws ~320 W

The iPhone 18 Pro number is Apple’s 50% figure on top of the A19 Pro’s 76.8 GB/s, since Apple doesn’t publish the absolute bandwidth. Against last year’s iPhone 17 Pro, my card is about 10× ahead instead of 6×.

  • A 20B model just fits on my card. At 4-bit, a 20B model is about 10–11 GB, so it sits entirely in VRAM. No flash streaming and no per-request routing needed. OpenAI’s gpt-oss-20b is the closest open equivalent: 21B total parameters, about 3.6B active per token, built to run in 16 GB. Since every expert stays in VRAM, it can switch experts every token, which the iPhone can’t afford.
  • My PC has the same problem one level up. Load a model bigger than 16 GB, like a 70B model at 4-bit (~40 GB), and the leftover layers spill into system RAM. Each token then waits on PCIe and the system RAM bus, and speed drops sharply. That’s the same pattern the iPhone has between RAM and flash, with a similar ratio between the fast and slow tiers. Tools like llama.cpp handle it by splitting layers between the GPU and CPU. Apple handles it by choosing a slice once per request.
  • My dictation setup is a smaller version of the same idea. I load Whisper large-v3 (~3 GB) into VRAM for each recording and free it afterward to keep room for ComfyUI. Apple does that with experts instead of whole models, on about 115 GB/s instead of 736.
  • Raw speed goes to the PC by a wide margin. About 6× the memory bandwidth plus far more compute means the card would generate several times faster on a comparable model. The iPhone wins on watts per token, on fitting in a pocket, and on not needing a 20B model’s worth of fast memory at all.

Takeaways for Building on Apple’s On-Device Models

If you’re building an app on Apple’s Foundation Models, in Swift or through a wrapper like expo-ai, here’s what I’d keep in mind after two rounds.

  1. Check capabilities at runtime instead of the iOS version. The same iOS release runs different models on different phones. iOS 27 on my 16 Pro has no image input. Read SystemLanguageModel.default.capabilities, or whatever your wrapper exposes, and give every feature a fallback. Test on real devices too, including the oldest one you support, because what works on one tier may not exist on another.
  2. Know exactly what the model sees. With the OCR tool, the model gets plain lines of text with the positions dropped. That works for a clean screenshot like a nutrition summary. It breaks on tables, because nothing says which row a number came from. Pick the input path that fits your data, or keep the positions yourself before handing the text over.
  3. Validate what the model extracts. Guided generation makes sure the answer has the right fields. It doesn’t make sure the numbers are right. My scale photo got a confident 250 in round one. Check anything you save against something you trust, like recent readings, and let the user confirm or edit before it lands.
  4. Handle tool failures. In round two the model called the OCR tool four times on unreadable text and hit the tool-call limit. Catch those errors, cap retries, and fall back to something useful, like manual entry.
  5. Keep requests small. The context window is 4,096 tokens on base devices, and the text from a busy screenshot eats a lot of it. Ask for one thing per request and trim what you send.

Where My App Goes From Here

I have three ways forward:

  • Model vision, which needs a 12 GB phone as far as I can tell. I’d have to buy a new phone.
  • OCR with positions, so table rows stay together. This one works on my phone today, and I could do some tinkering.
  • A cloud model, which defeats the point of this whole experiment.

The app is in good shape for the next phase. Clean screenshots work in OCR mode, which covers my macros from MyFitnessPal. Rejecting a scale reading more than about three pounds off my recent ones, plus manual entry as a fallback, should handle the scale. Although manual entry kind of defeats the point. But really, I want to combine all this data. Being able to pull in just one source isn’t that useful. The whole goal is to consolidate data siloed across different apps, which is what makes it tough to manage.

This post could just as well be called testing hardware. I assumed I’d get the new image input when I installed iOS 27. iOS 26 already had an image encoder that developers couldn’t reach, and iOS 27 opened it up, just not for my phone. I think it makes sense that Apple’s own features get these capabilities first, so I’d expect the public API to keep lagging a bit behind.

Making Big Models Fit

My 4080 Super has 16 GB of VRAM. For my AI animation project, I’ve run a 22-billion-parameter video model on it, LTX-2.5, and Wan 2.2, a mixture-of-experts model with about 27 GB of weights. Some of that came after I wrote that post, so you won’t find Wan or the quantized builds in it. Neither model fits as is. I got them running the same basic way Apple gets Core Advanced onto a phone: store the weights smaller, keep only part of the model in fast memory, and move the rest in and out.

Apple’s version is the library from earlier in this post. The full 20B model sits in flash, and each request loads a 1–4B slice into RAM. Mine is cruder, but it’s the same problem. The model is bigger than the fast memory, so the rest lives somewhere slower and only what’s needed gets pulled in.

The hardware got better too, but most of what made these fit came from people working out new tricks. I think that’s where a lot of the progress on local models is going to keep coming from. People keep finding ways to store weights in fewer bits, run less of the model at a time, and move weights around smarter, and those tricks stack. So I’d guess the models you can run at home will keep getting more capable faster than the hardware does.

The tricks cost something, though. I still had to run LTX at 768×512 instead of HD, and 2-bit weights lose some precision. Core Advanced still needs a 12 GB phone. The biggest models will probably stay in data centers, and what seems to be changing is how much you can do without them. For my animation work that matters a lot. A five-second clip costs me about half a cent in electricity at home, versus 2 on a cloud video API.

FAQ

Glossary