2.
Foundation Model Interactions
Written by Bill Morefield
In the first chapter, you developed a simple chat app that allows the user to send prompts and display responses from Foundation Models. This simple app lets you explore the basics of Foundation Models. Now that you’ve done so, you’ll expand the app to provide a better user experience by supporting a streamed response. Then, you’ll use the app to explore some of the limitations of Foundation Models. Finally, you’ll explore the LanguageModelSession and tokens.
Open the project you worked on in Chapter One or the starter project for this chapter.
Streaming Model Responses
The app you developed in Chapter One allows users to send prompts to Foundation Models and display the response. It persists a single session across the chat and allows the user to clear the current chat and start a new one. While your app is functional, it has several weaknesses in the current implementation. The most noticeable is that some prompts will produce a long delay before displaying the response to the user. To see this, run the app and enter a more complicated prompt.
Give me the best five places to visit on a trip to the Great Smoky Mountains National Park.
The response to this prompt will be lengthy. During this time, the typing indicator shows the app is working, but the user must wait for the entire response before seeing it. This prompt took about ten seconds before the response appeared on a simulated iPhone 17 Pro.
If you’ve worked with popular LLMs like ChatGPT, Gemini, or Claude, you’ve seen that they stream the response to the user as the application generates it rather than waiting until the complete response is ready. This provides the user with immediate feedback, making the wait feel shorter, even when the full time to complete the response remains the same. Foundation Models supports this streaming response capability.
Open ChatView.swift and find the sendPrompt() method. Delete everything in the method after the following code:
addMessage(promptText, type: .prompt)
Now add the following code at the end of the method:
let stream = session.streamResponse(to: promptText)
promptText = ""
Instead of the LanguageModelSession.Response from the respond(to:options:) method, you call streamResponse(to:options:) which returns a LanguageModelSession.ResponseStream<String>. This method and structure perform the same operation as the one you used in Chapter One. Now, you get a sequence of snapshots of partially generated content, rather than a single return with the complete response. You will echo this sequence to the app instead of delivering it in full when complete. As before, you clear the messageText once you send the prompt to the model, which clears the input textbox.
Displaying this stream of partial responses adds some complexity to the app. First, you’ll add a new Message type for a partial response. Open Message.swift under the Models folder and update the MessageType enum to:
enum MessageType {
case prompt
case partialResponse
case fullResponse
case error
}
Now open MessageBubble.swift to show this new type. Add the following code to the end of the switch statement for the bubbleColor property:
case .partialResponse:
return Color.gray.mix(with: .white, by: 0.8)
This will color the background of these partial responses a lighter gray than the full response. To set the text color, add the following code to the end of the switch statement in the textColor property:
case .partialResponse:
return Color.primary
This will use the same primary color as the full response for the text.
Now, return to ChatView.swift and add the following code to the end of the sendPrompt() method:
// 1
do {
// 2
for try await partialResponse in stream {
// 3
if messages.last?.type != .partialResponse {
// 4
addMessage(
partialResponse.content,
type: .partialResponse
)
} else {
// 5
messages[messages.count - 1].text = partialResponse.content
}
}
} catch {
// 6
addMessage(error.localizedDescription, type: .error)
}
Most of this code should look familiar. The general process of capturing the response remains the same. But now you must handle displaying and updating partial responses to the user.
- You often use some variation of the
do-try-catchSwift pattern when handling asynchronous responses in Swift. - The
streamreturned bystreamResponse(to:options:)is anAsyncSequence. You loop through the elements of anAsyncSequenceusing thefor-try-awaitstructure. Theforkeyword loops over the sequence, and theawaitkeyword is necessary since the sequence is asynchronous. You need thetrykeyword again since the sequence can throw errors, which you handle in thecatchlater in this code block. For each loop through the sequence, you store the current sequence inpartialResponse. - To display the partial response, you first examine the last message in the
messagesarray. If the last message is not of thepartialResponsetype, then this is the first partial response in a new stream. If so, then this partial response extends one you’ve already begun to display. -
partialResponseholds the current response. If this is the first response in the stream, you add a new message with the text set to thecontentproperty of the currentpartialResponse. Except for very short responses, more partial responses will follow. You note this by setting it as apartialResponsemessage. - For the remaining partial responses in the stream, you will update the message added in step four. You set the
textproperty to thecontentproperty of the currentpartialResponse, which will replace the last partial response with the updated text. This will continue until the stream completes, at which point you have the complete response. - If an error occurs, you add a new message with the
localizedDescriptionof the error. Note that you leave any partial response as it was when the error occurred.
Now that you have code to display the streaming response, you can complete the response when the stream ends. Add the following code to the end of the sendPrompt() method inside the for-try-await loop, just before the catch:
let lastIndex = messages.count - 1
withAnimation(.easeInOut) {
messages[lastIndex].type = .fullResponse
}
messages[lastIndex].timestamp = Date.now
This code will change the type of the message to .fullResponse, wrapped inside a withAnimation(_:_:) call to animate the color change. It also updates the timestamp to the current time.
Run the app and try the previous prompt. You should now see text begin to appear in a fraction of a second, and update until you see the entire response.
The result looks much better to the user as the text starts to appear after a few seconds. Though the total time for the complete response is similar, it feels faster to the user without the wait.
Why would you not use a streamed response? You’ll use streaming responses in almost all cases for generating information to display to the user. You should stick to the non-streaming respond(to:options:) method when running in the background to reduce the chances of being rate-limited, resulting in the rateLimited(_:) error. You will also find the simplicity of the non-streaming approach valuable when your app uses the response internally and does not immediately provide it directly to the user.
Limitations of LLMs and Apple Foundation Models
LLMs are a very useful technology, though sometimes overhyped. Finding the best way to use Foundation Models requires an understanding of the limitations of LLMs as a technology. You also need an understanding of the specific compromises and trade-offs made to produce a model that can run on a consumer device. You’ll use the app and some prompts that show the use cases where the model works well and where it can fail.
Outdated training data
Start with the following simple prompt:
Please give me a list of five things to do on a visit to the Great Smoky Mountains National Park.
The model will dutifully provide five activities. While yours may vary, they will likely all be reasonable things to do on a visit to this major national park.
However, the information provided here is not perfect. The first item on the list in the screenshot suggests hiking Laurel Falls Trail. As of April 2026, the trail has been closed since January 2025 for repairs, which are expected to last eighteen months. And while the second item, Clingman’s Dome, is a breathtaking view of the surrounding mountains and valleys, the park service changed the official name to Kuwohi in 2024.
Note: Remember, the responses you see may be different.
These mistakes both show the first weakness of LLMs. They only reflect the knowledge they were trained on. And they do not contain factual information beyond their training date. If you ask for factual information beyond that date, it will state this directly, responding that Apple Foundation Models does not have access to data beyond October 2023. This cutoff date will likely change in future versions of Foundation Models.
These hiking trail examples showcase a subtle way in which information not available until after the training date can yield poor results. Apple does provide a way to mitigate this limitation with tools, something you’ll explore in a later chapter. The main takeaway is that you should not count on recent factual information being present in the model. More importantly, you cannot count on the model knowing that it doesn’t know information from after this cutoff date.
This is not the same as a hallucination, which occurs when a model generates responses that are incorrect or nonsensical. Here, the information presented was correct at the time of training but has since become outdated.
Hallucinations
The second risk is better known as hallucinations. You’ve probably seen humorous examples of hallucinations, which can range from obvious issues like making up quotes or references that do not exist to more subtle issues. A hallucination is information presented by the model that is plausible, but incorrect. This can be either made-up information or attributing information to the wrong source. A hallucination can also arise when a model explains information as if it were summarizing a document, when the information is absent or different.
The danger in hallucinations comes in that they read or sound correct, but are not. It is important to check any information an LLM provides and ensure that your app can handle the situation. A hallucination in a bedtime story will probably have no serious effect. A mistaken measurement in a recipe can ruin the meal.
But sometimes the response breaks in clear ways. For a rather extreme example, at least as of iOS 26.4, enter the following prompt:
Give me a list of US states in reverse alphabetical order.
Enter the prompt, and you will see it go horribly wrong in one of several exciting ways. In the 26.4 versions of the Apple OS family, you’ll see some states left out and other states repeated. In the screenshot example below, Massachusetts appears nine times, with no sign of New Jersey, New Mexico, or New Hampshire. Variations on this prompt generally exhibit the same problem, and often it explodes into a long, repeating list of states until the context length fills up or the app crashes.
This seems like a fairly simple prompt to break Foundation Models. That’s the point, that you cannot assume anything will work without testing. You need to test your prompts before adding them to your apps. Some of the tradeoffs in making a model that runs on a device make hallucinations more common than in larger models. Testing is your best defense to reduce hallucinations in your app.
Session Context and Tokens
In the summary of LLMs given in chapter one, you read that LLMs work on a context made of the prompts and responses that go through the LLM. Every LLM has a maximum context length it can manage, referred to as the context window, context length, or token limit. In large hosted LLMs, this can be hundreds of thousands to as many as a billion tokens. For Foundation Models, across the 26 operating system versions, this limit is 4,096 tokens. And yes, the length is measured in tokens. Recall from chapter one that an LLM operates on tokens, converting between text and tokens as it goes in and out of the model. Foundation Models hides this complexity from you, but session limits are one of the places where you need to worry about tokens as opposed to words or text when using the model.
Sessions that near the token limit also face other challenges. Crossing this 4,096 token boundary will normally result in a hard failure with a LanguageModelSession.GenerationError.exceededContextWindowSize error. As the number of tokens in the context window size nears the limit, the time taken to generate a response will increase, as the model considers the entire context for each prompt. And until 26.4, you had no reliable way to track the session’s context length. The estimate that one token is roughly 0.75 words was about as close as you could get.
The first property you can access contains the model’s context window length. While that is constant now at 4.096 tokens, this property allows you to future-proof an app for future versions, which may extend the length or provide multiple models. Open ChatView.swift and add the following new property after the existing ones:
private var contextWindow = SystemLanguageModel.default.contextSize
This creates a new property and sets the value to SystemLanguageModel.default.contextSize. The SystemLanguageModel.default represents the primary, on-device model in Apple Foundation Models. The contextSize property contains the maximum context window size in tokens that this model supports. Now add the following code at the end of the VStack of this view, just after the MessageInputView and its modifier:
Text("Context Window: \(contextWindow) tokens.")
.font(.footnote)
Run the app, and you will see the maximum context window size added to the bottom of the view.
While most token-related methods require 26.4 or later operating systems, this property will be ported back to earlier operating system versions.
Counting Foundation Models Tokens
Knowing the maximum size is helpful because it lets you compare this to the sizes of your prompts and responses. Starting with the 26.4 release, you now have ways to read and count tokens much more easily. Add the following new property to the view:
@State private var contextWindowSize: Int?
This optional will hold the current context window size. Now add the following new method to the end of the ChatView struct:
// 1
private func updatedContextWindowUsed() async {
// 2
guard #available(iOS 26.4, macOS 26.4, *) else {
contextWindowSize = nil
return
}
// 3
contextWindowSize = try? await SystemLanguageModel.default.tokenCount(
for: session.transcript
)
}
Here’s how this method works:
- Since the calculation of the token length can take some time, all methods that do so are marked
async. Therefore, you need to mark the methodasync. - These token-related methods are only available in 26.4 operating systems or later. This guard statement ensures the app is running on iOS 26.4. If running on an earlier version, you could use another method to estimate the size, but this implementation sets the
contextWindowSizetoniland returns. - The
tokenCount(for:)method takes several possible parameters. Thesessionproperty, which holds your Foundation Models session, contains atranscriptproperty. Thetranscriptcontains a linear history of entries that form the interactions with a session. You will explore the Transcript more in Chapter Four. Passing this totokenCount(for:)will return the number of tokens in the session. You store this in thecontextWindowSizeproperty. If anything goes wrong during the method call, then thetry?will set the value tonil.
You now need to call this method whenever the session updates. The best place to update this is after you add messages to the chat. You already have a method that does this. Find the sendPrompt() method and add the following code to the end of the method after the do-try-catch block:
await updatedContextWindowUsed()
Now, whenever you add new messages to the chat, you update the context window size. To show this to the user, update the Text view at the end of the VStack to:
if let tokenCount = contextWindowSize {
Text("Context Window: \(tokenCount)/\(contextWindow) tokens.")
.font(.footnote)
} else {
Text("Context Window: \(contextWindow) tokens.")
.font(.footnote)
}
This code attempts to unwrap contextWindowSize. If successful, the view will display that value along with the maximum context window length. Otherwise, you show the maximum context window size. Build and run to see this in action.
Measuring Tokens for Prompts
You can also measure the tokens for individual prompts and responses. You’ll first update the Message struct to hold a token count. Open Message.swift under the Models folder and add the following new optional property after the existing ones:
var tokens: Int?
This adds an optional Int to hold the message’s token count. Update the initializer to:
init(
id: UUID,
text: String,
type: MessageType,
timestamp: Date,
tokens: Int? = nil
) {
self.id = id
self.text = text
self.type = type
self.timestamp = timestamp
self.tokens = tokens
}
This adds a parameter for the new tokens property to the initializer with a default of nil if not provided. Back in ChatView.swift, add a new method after the existing ones with the following code:
private func tokenCount(for text: String) async -> Int? {
guard #available(iOS 26.4, macOS 26.4, *) else { return nil }
return try? await SystemLanguageModel.default.tokenCount(for: Prompt(text))
}
This short method first ensures the code is running on a version that supports the new token count methods. If not, the code returns a nil. If so, then it calls a tokenCount(for:) method on the default Foundation Model. Again, note that this method is asynchronous, marked async, and uses the try? keyword. If the call throws any exception, the call returns a nil value.
To add the tokens to the message, find addMessage(_:type:animate:). Since your tokenCount(for:) method is asynchronous, you need to update this method to work with asynchronous methods. A simple way to do this for this method is to wrap the entire method body inside a Task, making the existing code its closure. Once you’ve added the Task, add the following code to the top of the new Task’s closure:
var tokens: Int?
if type == .prompt || type == .fullResponse {
tokens = await tokenCount(for: message)
} else {
tokens = nil
}
You first declare an optional Int named tokens that will hold the number of tokens. The only two message types you want to show tokens for are prompt and fullResponse. Errors aren’t part of the session, so there’s no purpose in showing tokens for them. We also ignore partialResponse replies for two reasons. First, the number of tokens for a partial response provides no useful information since the response is incomplete and will change. Second, the time it takes for the asynchronous tokenCount(for:) method to complete will probably be longer than the time before the partial response is replaced. You set tokens to nil for the types where there is nothing to show.
Now update the assignment of newMessage to:
let newMessage = Message(
id: UUID(),
text: message,
type: type,
timestamp: Date(),
tokens: tokens
)
This will add the tokens to the method. Next, find sendPrompt() and look for the last lines of the do closure that reads messages[lastIndex].timestamp = Date.now. Add this new line after it:
messages[lastIndex].tokens = await tokenCount(for: messages[lastIndex].text)
This line retrieves the token count for the final response and sets the token property of the message to it.
Finally, you’ll add code to show the tokens along with the current timestamp. Open MessageBubble.swift. Find the closing Text view that references message.timestamp that reads Text(message.timestamp, style: .time). Replace this line with the following, leaving the font, foregroundColor, and padding as is:
HStack {
if let tokens = message.tokens {
Text("\(tokens) tokens")
}
Text(message.timestamp, style: .time)
}
This code wraps the existing Text view in an HStack view, moving the font, foregroundColor, and padding modifiers from the Text to the new HStack. Before the time, the code attempts to unwrap the message.tokens property. If successful, it will display that unwrapped value in the view.
Run the app and enter a prompt to see the results:
You might notice that the sum of the tokens counts for the prompt and response doesn’t match the tokens shown for the entire transcript. This is because you are calculating the number of tokens in the text. This may not be the same as the number of tokens this text uses in the context of the session. A more accurate value would require finding the entry in the Transcript and calculating tokens from it. For most apps, the difference will not be enough to matter.
You can trigger this error by having the session’s length exceed the token count. In a real app, you would need to handle this depending on your use case. You might create a new, empty session and start over. You could also summarize the current session and feed it into a new session to retain some context. You’ll look at some of these options in a later chapter.
Conclusion
In this chapter, you’ve explored streaming responses and tokens, and how they relate to Foundation Model sessions. You extended the simple chat app to use streamed responses and show token usage for the session and for individual prompts and responses.
Now that you’ve learned the basics of Foundation Models, you’ll look at tuning and guiding responses in the next chapter.
Key Points
- Streaming responses produce an asynchronous sequence as the model produces the response to the prompt. Showing this will make your app feel more responsive.
- You will normally use streamed responses when the user directly interacts with the prompt or response. Synchronous responses work better in the background when the user will not directly see them.
- Up through 26.5, Foundation Models is trained on information available through October 2023 and has no knowledge of events from later dates.
- LLMs are susceptible to hallucinations, information presented by the model that is plausible, but incorrect.
- LLMs will also sometimes not get the prompt right. Sometimes you can correct this by adjusting your prompt.
- Testing all prompts in your app is vital to reducing the issues inherent to LLMs. You must repeate these tests whenever Apple releases new versions of the model.
- Every LLM has a maximum length of tokens it can contain. For Foundation Models, this is 4.096 tokens.
- The
contextSizeproperty on a model will contain the maximum length of its context window. Apple introduced this in 26.4, but backported it to earlier 26 OS versions. - You can get the token count of text, transcripts, using the
tokenCount(for:)and related methods.