1.
Introduction to Foundation Models
Written by Bill Morefield
Machine learning and other artificial intelligence systems have gone from curiosities to useful tools in the last few years. While their use cases are often overhyped, the fact that they can provide clear value to your app in the right situations is clear. You can now complete tasks on a device that fits in your hand that were previously difficult or impossible, even on enterprise equipment. Only a few of these technologies have garnered the hype and controversy as Large Language Models (LLMs). A traditional LLM requires a massive amount of computational power, memory, and resources to run. Well-funded startups were the one ones able to train and run these models and deploy huge amounts of memory, storage, and computing power.
To address these high system requirements, many programmers have explored deploying local models. These models are optimized and simplified to run on the devices and equipment of everyday users. Starting with iOS 26, iPadOS 26, macOS 26, and other version 26 operating systems, Apple is providing its own local model, optimized for use in apps, called Apple Foundation Models. Apple Foundation Models are Apple’s on-device AI models, designed to protect privacy while helping with tasks like writing text, summarizing information, and organizing data on supported devices. Because all data remains on the device, you don’t need an internet connection, there is less latency, and you avoid the privacy risks that come when sending data to third-party services. This book will explore the use of Apple Foundation Models in your apps.
What is Apple Foundation Models?
It’s worth starting with the most basic question: What is Apple Foundation Models Framework? The short answer is that it is a large language model (LLM) that Apple has optimized to run locally on end-user devices, such as laptops, desktops, and mobile devices. Traditional LLMs operate in data centers equipped with high-powered GPUs, which need a lot of memory and power. Bringing that functionality to an end-user device requires significant changes to the model. In Apple’s case, the two most important changes to produce Foundation Models are reducing the number of parameters and quantizing the values that form the model. You’ll learn more about how they did that later.
In this chapter, you will develop a chat-style app that interacts with Foundation Models to explore the possibilities and limitations of this framework. A chat app isn’t a great use case for Foundation Models due to the small size of the on-device model, but it provides a well-understood app to explore integrating Apple Foundation Models into a SwiftUI app. The immediate feedback will also make it easier to explore the model’s use and limitations.
Using Foundation Models
Open the starter app in Xcode 26 or later. Run the app, and you will see the starter implements a simple chat-style interface. The textbox at the bottom of the view provides the user a place to enter text and “send” it. Right now, the chat will echo any text entered.
Note: When running apps that use Foundation Models in the simulator, the simulator uses the underlying device’s Apple Intelligence. That means you should run on macOS at least the same version as Xcode and the simulator you are using. This includes beta versions. You cannot run a simulator with iOS 2.64 beta on macOS 26.3 and simulate Foundation Models. If the versions don’t line up, using Foundation Models will return an error.
You’ll now update this app to use Apple Foundation Models. The text you send to the model is a prompt. The model will then provide a response, which the app will display to the user.
Before you try to use Foundation Models, ensure that it is available on the device. The device must support Apple Intelligence, and the user must turn it on for their device. If either of those is not done, then your app will need to disable or work around model features. So your first step is to ensure that Foundation Models are available. Currently, the app only shows the ChatView, but you will instead display a message if the device does not support Foundation Models.
Open ContentView.swift and add the following code as the last import:
import FoundationModels
To use Foundation Models in a class or view, you must first include the framework. Now add the following property to the view:
private let model = SystemLanguageModel.default
SystemLanguageModel refers to the on-device text foundation model. The default property accesses the base version of the model. You will use this to verify the status of Foundation Models. Replace the body of the view with:
switch model.availability {
case .available:
ChatView()
case .unavailable(let reason):
ModelUnavailableView(reason: reason)
}
ModelUnavailableView doesn’t exist yet, but you’ll take care of that next. The model.availability property reflects the state of Foundation Models on the device. It will either be available, meaning your app can access Foundation Models, or unavailable, which also provides a reason explaining why Foundation Models is not accessible. When available, you show the existing ChatView as before. But when Foundation Models is unavailable, you will display ModelUnavailableView to inform the user.
Create a new SwiftUI view named ModelUnavailableView in the Chat Views folder. Add the following to the imports at the top of the new view:
import FoundationModels
Again, you need this import in any code that interacts with Foundation Models. Next, add the following reason property to pass into the view:
var reason: SystemLanguageModel.Availability.UnavailableReason
This property holds an enumerable that explains why the model is unavailable. Your app can use this to display an appropriate message. Change the body of the view to:
Image(systemName: "apple.intelligence")
.font(.largeTitle)
switch reason {
case .deviceNotEligible:
Text("Apple Intelligence is not available on this device.")
case .appleIntelligenceNotEnabled:
Text("Apple Intelligence is available, but not enabled on this device.")
case .modelNotReady:
Text("The model isn't ready. This is usually because it is still downloading.")
@unknown default:
Text("An unknown error prevents Apple Intelligence from working.")
}
This will display the Apple Intelligence SF Symbol along with a user-friendly text message for the most common reasons. You also provide a generic error for other cases, using @unknown default to future-proof against new enum values. This will prepare your app for any future changes to the framework. Now update the preview to:
ModelUnavailableView(reason: .appleIntelligenceNotEnabled)
This provides a reason for the preview. Viewing the Canvas will now show this default view.
Now, run your app. If your device meets the requirements described earlier, you should still be able to see the chat app. If your device doesn’t support Apple Intelligence, you will see the informational view to that effect, along with the reason. Of course, when testing your app, you will want to ensure the user either gets a successful fallback or an appropriate informational message for these error states. To test this, you can use the scheme option in XCode.
Note: If this is the first time you are using Apple Foundation Models, it could take 15 to 60 minutes for the model to download on the device. Make sure the device has a connection to the internet while it’s downloading.
Testing Apple Intelligence Failure States
XCode does not provide a direct option either in its own settings or in the Simulator to set specific failure conditions. You can accomplish this using schemes. In XCode, select Product ▸ Scheme ▸ Edit Scheme….
Change to the Options tab. Scroll near the bottom of the list and you will see an option Simulated Foundation Models Availability with a dropdown providing five choices. When the default Off is selected, there is no change to the state of the simulator. The other options will produce the listed error condition for Apple Intelligence regardless of the device’s settings or capabilities. For now, change it to Device Not Eligible.
Click Close and run the app again. You will see that the app shows that Apple Intelligence is not available on this device.
With this, you can verify that the fallback processes and the error messages in your app work for the common cases where your app will not have access to Foundation Models. Make sure to go back and change the Scheme to Off before continuing in the chapter.
Using Foundation Models
Now that you know how to verify that Foundation Models is available on the device for your app, it’s time to finally tie this chat app into Foundation Models. Open ChatView.swift. First, import Foundation Models by adding the following import after the existing one.
import FoundationModels
Now find the sendPrompt method. Interactions with LLMs consist of a prompt sent to the model and a response from the model. In this app, the text entered by the user will be the prompt. You will then take the response from the model and add it as a “reply” to the messages list.
Replace the current method contents with:
// 1
guard !promptText.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty else { return }
// 2
addMessage(promptText, type: .prompt)
// 3
let session = LanguageModelSession()
do {
// 4
let modelResponse = try await session.respond(to: promptText)
promptText = ""
// 5
addMessage(modelResponse.content, type: .fullResponse)
} catch {
// 6
let errorResponse = "An error occurred while processing your message. \(error.localizedDescription)"
addMessage(errorResponse, type: .error)
}
The most valuable benefit of Foundation Models comes from the simplicity of using it. That code provides a basic, but complete implementation of sending a prompt to the model and getting the response:
- You ensure there is useful text in the
promptTextand return without doing anything if there is no prompt to process. - You add the prompt to the list of messages using the
addMessage(_:type:animate:)method, passing the text of the prompt along with aMessageTypeenumeration indicating that this is a user prompt. - You use a
LanguageModelSessionto interact with Foundation Models. This represents a single session of interactions with the language model. You’ll learn more about what a session means throughout this book. - The
respond(to:options:)method sends a prompt to theLanguageModelSessionthat you created with a string prompt and returns aLanguageModelSession.Response. Therespond(to:options:)method generates the entire response and returns it when ready, which can take some time for complex or long prompts. Apple therefore made the calls to it asynchronous. The method will wait until the call returns before continuing. Because of this, you need toawaitits completion before continuing. Also note that this method is marked asasyncto accommodate this need. If the model returns a valid response, you clear out the user’s input text. - You use the
contentproperty of the result to get the text reply. You then add the response text to the message list using thefullResponsetype. Note that LLMs, including Foundation Models, often return text in Markdown, a simple markup language that allows styling while still written with plain text. TheMessageBubbleview already supports Markdown text and displays this content correctly. - If anything goes wrong, you put the error message into
errorResponseand add it to the messages, passing theerrorenumeration so it is formatted as such. You will learn more about error handling later in this chapter.
Run your app. Enter a simple prompt and tap the send button. After a pause of several seconds to a minute, you will get a response from the model.
Take a moment to appreciate that you have set up and used an LLM in just two lines of code, excluding the code required for displaying text and handling errors. That provides a convenience that hasn’t been available when using LLMs before.
Note: Do not worry if your replies in this and the following chapters don’t precisely match the ones shown in the lesson. LLMs are probabilistic by nature, meaning there is some randomness implicit in the system. You can adjust this and even eliminate this randomness using options, which you’ll learn about later in the book. For now, as long as the responses are fairly reasonable, you are probably seeing the correct behavior.
What is an LLM?
Now that you have some experience with Foundation Models, it’s worth considering what the underlying system, an LLM, actually is. Just the detailed discussion of how LLMs work could fill an entire book. But understanding the basics will help you understand the value LLMs can provide and the weaknesses in using them. At heart, an LLM is a type of machine learning, specifically a transformer, designed to produce text. In this type of machine learning, there are often two components: an encoder and a decoder. Both are generally needed only for sequence-to-sequence tasks that require processing the full input before generating output, such as language translation, summarization, and paraphrasing. Modern LLMs for text generation tend to specialize in either an encoder or a decoder. Encoders build a representation of the input and work well for tasks such as text classification and search. Decoders produce better results to create open-ended text. Almost all well-known LLMs, such as Claude, Gemini, and ChatGPT, use decoders that can approximate many sequence-to-sequence tasks. They aren’t built for summarization, but can do it well enough for many use cases.
A decoder is created by training, a process that requires vast amounts of text. Some of the controversy around LLMs comes from the acquisition of these large amounts of text. The ethics of using copyrighted text without permission to train the model are debatable, and courts worldwide are still determining the legality. The training process can be tuned or adjusted to produce a decoder that generates text more closely matching the desired output. Roughly speaking, the larger the model, measured by parameter count, the better it will perform on a given task, all other things being equal. Each parameter represents a single value inside the machine learning model. These can be stored using different data structures. The most common is the equivalent of the Swift Float type, which represents a number in a 32-bit format (often referred to in machine learning as fp32). This means each parameter takes four bytes. Large models require substantial memory to hold all these parameters.
How does the resulting decoder create text? The decoder does text prediction. Take the following start: I went for a. It is a valid English phrase, but it is incomplete. Any word could follow a, but only a small number produce a valid sentence. Only some of those valid sentences make sense. A person knows that the next word should be something a person can do. But machine learning doesn’t reason about what people can do. Instead, a machine learning system will predict the next word by choosing the most likely word to follow this text pattern. The most common next words in this case might be walk, run, swim, and ride. The context of the words before the one to be added influences the probability. If the previous text discussed a river, then swim becomes more likely. If the prior text mentioned horses or cars, ride becomes more likely. In reality, a system that always chooses the most likely word creates repetitive text. By default, the LLM selects the next word based on relative probabilities. This is why LLMs are non-deterministic by default. That is, providing the same input to an LLM can and usually will produce different outputs. Most models provide a parameter that allows adjusting how this choice is made, and many can be made more deterministic when desired, including Foundation Models.
In practice, modern LLMs act on tokens, not words, because text includes elements beyond words. Many tokens do represent common words one-to-one, but less common words are broken into chunks of a few characters that make up the token. Take the earlier example completed as I went for a run.. This sentence of 17 characters can be broken into six tokens: [I][ went][ for][ a][ run][.]. You can see that in this case, each word is a token, including the final punctuation. But take a sentence with less familiar words: I read Proust and Chaucer in class. would be broken into [I][ read][ P][rou][st][ and][ Chau][cer][ in][ class][.]. Note that the proper nouns of the author names break into two or three tokens while the more common words remain single tokens. A good rule of thumb is that a token represents about four characters and a bit less than a full word, which averages five characters in English. Other languages will have different relationships between text and tokens. You’ll explore tokens more later in this book, but you can use LLMs without worrying about them in most cases.
When you feed your prompt into the model, it is converted into tokens, and then sent to the LLM as input. The LLM then predicts text based on this prompt, using the data it was trained on and previously provided information, and returns a response. The accumulation of prompts and responses generates a context. The context contains all the information the LLM can reference, in addition to its training data, when processing a prompt. All models, though, have a maximum context length they can work with. Even for models with long context lengths, the quality of their output degrades as context length increases.
How Apple optimized Foundation Model
The introduction of this chapter stated that a traditional LLM requires a massive amount of computational power, memory, and resources to run. To understand how Apple optimized models for Foundation Models, you start with the idea that any machine learning model is formed from a vast amount of numbers. In LLMs, these numbers are called parameters. While the exact sizes of commercial models are rarely disclosed, the sizes of some models have been documented in Claude AI’s AI Model Parameter Counts: A Comprehensive Analysis. Recent models often contain hundreds of billions of parameters. Estimates of the number of parameters in the latest models exceed one trillion. As each parameter consists of a number, that number must be represented in a format that computers understand. The most common is a 32-bit floating point, abbreviated as fp32. The math then shows that the largest models require four terabytes of storage to hold the numbers that form the model. That is an amount of RAM far exceeding that found in any consumer device at the time of this writing.
The first step to reduce this number to something that will work on an end-user device is to reduce the number of parameters. The details on how to do this would be a course on its own, but with careful techniques, you can reduce the number of parameters while still having a helpful model, though with limitations from the source model. The method to achieve this will vary depending on the original LLM, the intended use case of the smaller LLM, and the acceptable performance requirements. The result will be a less capable but smaller model.
The second step is to reduce the size of each parameter by using a format that uses fewer than four bytes of an fp32. This technique, known as quantization, reduces these numbers to a lower precision format, in other words, fewer decimal points. This also has the advantage of speeding up model processing and reducing power demands, which is particularly valuable on a mobile device with a finite battery. When done carefully, quantization produces a significantly smaller model with minimal impact on output quality.
Apple Foundation Models contain approximately 3 billion parameters and are quantized so that each parameter occupies 2 bits. This lets the entire model fit in eight gigabytes, enough to fit on the newest Apple devices at the time of the initial 2025 release. This smaller size does create trade-offs, ones that Apple admits and documents. Apple tuned the model for text generation tasks, including summarization, entity extraction, text understanding, text refinement, dialog for games, and creative content generation. You can find more information on the training of Apple’s models in Introducing Apple’s On-Device and Server Foundation Models.
Handling Model Delays
While this basic implementation shows how little code you need to work with Foundation Models, it has several weaknesses. The most glaring is that you create a new LanguageModelSession for each prompt. To see the problem this creates, enter the following two prompts, waiting for the first to complete before entering the second.
Give me five popular fruits.
and
Which of these are commonly available in the United States in the summer?
The wording of the second response will vary, but it will generally show no idea of the fruits you asked about in the first prompt.
When you create a new session for each prompt, each exists as a stand-alone interaction. When you entered the second prompt, the new session knew nothing about the first prompt or response. To fix that, you should create a single session and send each prompt to it. Doing so is simple.
Open ChatView.swift and add the following new property to the view:
@State private var session = LanguageModelSession()
This creates a view property that holds a session. As long as you reuse this session, the session will keep an awareness of all prompts and responses. Now, find comment three in the sendPrompt method and delete the let session = LanguageModelSession() line. The method will now use the view’s session property, which will persist across multiple prompts.
Type the same two prompts asking about fruit again. This time, the second response will build upon the context created by the first prompt and response and answer reflecting the five fruits listed in the first response. Using a single session lets you provide multiple prompts and responses that become known information for later prompts.
Reusing a single LanguageModelSession introduced a few challenges. Because generating a response takes time, a shared session can encounter errors if you send a new request before the previous one completes. To see this in action, enter a prompt in the app and hit the send button twice in rapid succession. You will see this causes an error.
This well-written error contains the solution. To prevent this, add the following modifier to the MessageInputView view:
.disabled(session.isResponding)
This change disables the input while the session responds, so the user cannot send a second message until the first response is complete. You can also now use this property to provide a visual indicator when the model is working.
At the end of the ScrollView, add the following code:
if session.isResponding {
TypingIndicator()
.transition(.scale)
}
This will display the typing indicator when the session is responding to a prompt, giving the user a visual indicator that the app is working.
To this point, there has been no good way to clear a chat so the user can start over. To fix that, first add the following new method after sendPrompt():
private func resetChatHistory() {
messages = []
session = LanguageModelSession()
}
This first sets messages to a new empty array, clearing the existing messages. It then sets session to a new LanguageModelSession. As you saw earlier, this gives the app a fresh, clear session to work with. To give the user a way to invoke this, you will add a toolbar to the app. Add the following code to the end of the current properties:
@State private var confirmClear: Bool = false
Now add a new property to hold a toolbar that you will use to show the option. You will add more options to this toolbar throughout this book. Add the following code before the body of the view:
@ToolbarContentBuilder private var appToolbar: some ToolbarContent {
ToolbarSpacer(.flexible, placement: .bottomBar)
ToolbarItem(placement: .bottomBar) {
Button("Clear", systemImage: "xmark.circle.fill") {
confirmClear = true
}
.tint(.red)
.confirmationDialog(
"Are you sure you want to delete the chat history?",
isPresented: $confirmClear
) {
Button("Delete Chat History", role: .destructive) {
resetChatHistory()
}
}
}
}
This toolbar contains a single button that, when tapped, displays a confirmation dialog to the user. When the user taps the Delete Chat History button, the app calls the resetChatHistory() method, clearing the chat. To add this toolbar to the view, add the following code to the end of the VStack, just after the navigationBarTitleDisplayMode method:
.toolbar {
appToolbar
}
Run the app to confirm this works. Enter a few prompts, and then tap the red icon at the bottom of the window. Tap the Delete Chat History button. The existing messages should disappear, and prompts referencing them no longer work.
Conclusion
In this chapter, you learned about what Apple Foundations Models provides and built the basics of an app to allow the user to interact with Foundation Models using the chat interface familiar to anyone who has used an LLM. Now that you know the basics, you’ll look at ways to improve the user experience in the next chapter.
Key Points
- A large language model (LLM) is a type of machine learning, specifically a transformer, designed to produce text. There are often two components: an encoder and a decoder.
- A traditional LLM requires a massive amount of computational power, memory, and resources to run.
- Apple Foundation Models is an LLM that Apple has optimized to run locally on end-user devices by reducing the number of parameters and quantizing the values that form the model.
-
SystemLanguageModelis the on-device text foundation model. - You can test different failure scenarios using schemes.
- Interactions with LLMs consist of a prompt sent to the model and a response from the model.
- Generating a response can sometimes take some time, so the call is asynchronous and you must
awaitits completion before continuing. - Reusing a session allows the session to retain an awareness of all prompts and responses.