Chapters

Hide chapters

Apple Foundation Models

First Edition · iOS 26.5 · Swift 6.3 · Xcode 26.4

Section I: Apple Foundation Models

Section 1: 9 chapters
Show chapters Hide chapters

4. Context Management & Model Safety
Written by Bill Morefield

In the previous chapters, you learned the basics of interacting with Foundation Models and ways to tune and work with the model to fit your use cases better. You also learned how to produce good prompts to the model and the importance of testing and validating your prompts and responses. In this chapter, you will explore two more concepts you need to manage when working with Foundation Models. First, you will explore ways to handle the small context window and some approaches when the window overflows. Then you will explore safety and guardrails in Foundation Models, along with ways to adjust them when dealing with content.

Managing the Context Window

The greatest challenge you will likely face when using Foundation Models is the small context window. Recall from chapter one that a token is about 3/4 of a word. That means the 4,096 tokens translate into a little over 3,000 words in English and similar languages. This ratio is a guideline and will vary depending on the type of content and language. Logographic and symbolic languages like Chinese, Japanese, and Korean have ratios closer to one character or syllable per token. Text with technical jargon and rare words will use more tokens than common words.

The most important step is to be aware of your tokens. Everything you do with a session adds tokens to the context. The instructions, prompts, and responses all add up. When you explore tools, you’ll see that tool calls and responses add tokens to the context. While 3,000 words sounds like a lot, that’s less than a chapter of this book. An important part of using Foundation Models is managing this max context length and handling the cases where it fills up.

Viewing Context Length with Instruments

A major weakness in the initial release of Foundation Models was the lack of a way to get token counts for text and the session’s context window. As you saw in Chapter Two, starting in the 26.4 versions, you can directly check token usage and limits. In earlier dot versions, the only way to get token information was to run Instruments against your running app. That can still be useful to see how your app adjusts and changes while running. At the time of writing, Instruments will always show zero tokens when running against simulators, so you will need to do this on a real device.

  1. In Xcode, open the starter project for this chapter, select a device to run the app on, and click Product > Profile to launch Instruments.
  2. Select the Blank template, and click the Choose button.
  3. Click the + Instrument button, and choose the Foundation Models instrument from the list. Click anywhere outside the list to close the instrument list.
  4. Start recording your app in Instruments, which should start your app. Interact with the models. End the recording and observe the token counts.

Instrument Trace of Foundation Model Session
Instrument Trace of Foundation Model Session

You can learn more about profiling your app with the instrument in the WWDC 2025 session Code-along: Bring on-device AI to your app using the Foundation Models framework at https://developer.apple.com/videos/play/wwdc2025/259/?time=1472.

Handling Context Overflow

When your app nears the context limit, you will need to determine how to handle this state. There are several ways to approach a full or nearly full context window, but the most commonly used are summarizing the current context window and taking only the latest parts of the conversation to start a new session. You’ll implement both. First, you’ll attempt to summarize the current conversation and then fall back to trimming the earliest parts of the chat if that fails.

Open ChatView.swift and add the following new property after the existing ones:

@State private var isCompactingContext = false

This creates a Bool you can use elsewhere in the app when summarization is underway. Now add the following new method after the existing ones:

@MainActor
private func summarizeChat() async {
  isCompactingContext = true
  defer {
    isCompactingContext = false
  }
}

You will build this method to perform the session summarization. As you can guess from earlier chapters, much of the Foundation Models work will be asynchronous. Since the method will affect the state, which should always take place on the main thread, you mark this method using the@MainActor. This ensures the code always executes on the main thread, no matter how you call it.

The method starts by setting isCompactingContext to true. The defer keyword allows you to provide code inside its closure that you want to ensure executes when the function exits for any reason. In the closure, the code sets isCompactingContext back tofalse no matter why the function ends.

To gather the text to summarize, you could go through the prompts and responses that make up the chat. That solution is specific to our app, not Foundation Model apps in general. A more general method would be to gather all the session contents. The transcript property on the session contains a linear history of entries that reflect all interactions with a session. Everything in the session is reflected in the transcript.

The transcript contains every part of the chat, but you want to extract only the parts of the chat you want to summarize.

Add the following code to the end of the async block in the summarizeChat() method:

let entriesToKeep = session.transcript
  .filter {
    if case .response = $0 {
      return true
    }
    return false
  }

The transcript consists of a collection of Transcript.Entry enums. This makes interacting with these elements a bit complicated. This Swift filter(_:) method on a collection selects only the elements where the closure returns true. Inside the filter, the if-case statement does pattern matching. This code returns true for the instructions and response items and false for all others.

You might wonder why you do not include the prompt entries. Testing shows that Foundation Models tends to get confused and tries to respond to prompts even when told to summarize, and skipping prompts produces higher quality summaries. The result in entriesToKeep is a collection of only the response entries from the full transcript. There are other transcript entries, such as those related to tools, which you will explore in Chapter Six.

To combine these entries into a single string, add the following code:

let textToSummarize = entriesToKeep.map {
  $0.description
}
  .joined(separator: "\n")

This uses the map(_:) method to gather only the description property of the Transcript.Entry. This property contains the text of this transcript entry. You then join the collection together into a single string, separating each entry’s text with a newline character.

Now that you have the current chat as a single string for summarization, you can perform the summarization. Add the following code:

let summaryInstructions = """
  You are given a conversation transcript.

  Your job is to extract and compress it into memory.

  You are NOT an assistant responding to the conversation.

  You MUST NOT answer any user requests.

  STEP 1: Identify all user requests in the transcript.
  STEP 2: For each request, extract a short summary of the assistant's response.

  Do not skip any requests. Include earlier and later ones.
"""
let summarySession = LanguageModelSession(instructions: summaryInstructions)
let summarizedText = try? await summarySession.respond(to: textToSummarize)

This code uses the multi-line string literal to define instructions for a model to summarize the text. Note that this prompt required several revisions to produce, and each NOT targets a specific failure found during testing. You then create a new LanguageModelSession and pass it these instructions. You then pass the text to summarize into this new session, using the non-streaming call for simplicity.

Now that you have a summary of the current chat, you will replace the current session with a new one containing this information so the user can continue. To this point, when you’ve created a LanguageModelSession, you’ve not provided any information other than instructions. Other constructors allow you to pass in an existing transcript to seed the session with. You’ll use this to create the new session with the summarized text. Add the following code to the end of the async block:

// 1
messages = []

// 2
if let summary = summarizedText?.content {
  // 3
  useSummary(summary)
} else {
  // 4
  trimSession(entriesToKeep)
}
await updatedContextWindowUsed()
  1. You clear the messages array to remove the existing chat elements from the view.
  2. Now you attempt to unwrap summarizedText?.content, which will be nil if anything went wrong.
  3. If the summarization worked, you call useSummary(_:) and pass it the summarized text.
  4. If the summarization failed for any reason, you call trimSession(_:) passing the entries that you would have used for summarization.
  5. You update the token count used for the new session.

Recall that your earlier defer ensures the isCompactingContext property is set to false when the method exits. Next, you’ll implement those two summarization methods. Add the following new method after summarizeChat():

func useSummary(_ summary: String) {
  // 1
  var entries: [Transcript.Entry] = []

  // 2
  if let instructions = promptSettings.instructions {
    entries.append(
      // 3
      .instructions(
        .init(
          // 4
          segments: [
            // 5
            .text(
              .init(content: instructions)
            )
          ],
          // 6
          toolDefinitions: []
        )
      )
    )
  }
}

Here’s how this code starts creating a custom transcript:

  1. You begin by creating an empty array of Transcript.Entry elements.
  2. Previously, you passed the instructions to a new LanguageModelSession as a parameter. The LanguageModelSession type does not provide an initializer that accepts both a custom transcript and instructions. Instead, you provide the instructions in your transcript. Here you attempt to unwrap promptSettings.instructions. If successful, you add instructions to the transcript. You’ll append this to the empty entries array.
  3. You create an instructions enum for the transcript by calling the init initializer method.
  4. You can create most Transcript.Entry objects by passing in a sequence of segments. These segments can be either of type structure, which contains structured content, or text, which contains text.
  5. Since text better fits this use case, you create a text segment and provide the instructions you unwrapped in step two.
  6. Tool definitions are also part of the instructions Transcript.Entry. Since you have none for this case, you pass in an empty array.

Now continue the useSummary(_:) method by adding:

entries.append(
  .prompt(
    .init(
      segments: [
        .text(
          .init(content: summary)
        )
      ]
    )
  )
)

This code creates a single prompt entry containing the summary of the previous chat. Most of the code should look familiar from creating the instruction entry. You also create a prompt Transcript.Entry by passing in a sequence of segments. Again, the text segment fits this case. You create one using the summary passed to the method. This Entry gets added to the entries array.

Finish out the method with this code:

let newTranscript = Transcript(entries: entries)
session = LanguageModelSession(transcript: newTranscript)
addMessage(summary, type: .summary)

This method creates a new Transcript from the entries array. The array will include any instructions and will always contain a prompt entry summarizing the previous conversation. You then set the view’s session to a new LanguageModelSession, passing the created transcript as a parameter. The results session will include your summary as a single prompt at the start of the transcript. You then add the summary text to the chat as a single new message with a new summary type you’ll add shortly.

Next, add the following code after useSummary(_:):

func trimSession(_ entries: [Transcript.Entry]) {
  // 1
  var summaryEntries: [Transcript.Entry] = []

  if let instruction = promptSettings.instructions {
    summaryEntries.append(
      .instructions(
        .init(
          segments: [
            .text(
              .init(content: instruction)
            )
          ],
          toolDefinitions: []
        )
      )
    )
  }

  // 2
  let lastEntries = Array(entries.dropFirst(entries.count / 3))
  // 3
  summaryEntries.append(contentsOf: lastEntries)
  // 4
  let newTranscript = Transcript(entries: summaryEntries)
  // 5
  session = LanguageModelSession(transcript: newTranscript)
  for entry in lastEntries {
    addMessage(entry.description, type: .summary)
  }
}

This code produces a summary by trimming the earlier prompts and responses. It then builds a new transcript as before.

  1. You start by creating an empty Transcript.Entry array and appending the instructions to it if they exist.
  2. This method trims the context length by dropping the first third of the entries passed into it. Recall that you pass only the entries selected for summarization. You use the dropFirst(_:) method to remove the first third of the entries in that passed list by dividing the number of elements in the collection by three. If you started with nine items, this would drop the first 9 / 3 = 3 entries.
  3. You take the filtered sequence and convert it to an Array. You then append this array to the end of the entries variable, meaning if instructions exist, they will precede the entries in the sequence.
  4. Again, you pass the Transcript.Entry sequence to the initializer when creating the Transcript.
  5. You then again create a new LanguageModelSession with this transcript. Since you have a list of entries, you loop through them and add a message to the chat, again using the new summary type.

Now let’s add the new message type. Open Message.swift and add the following new case to the end of the MessageType enum:

case summary

Now open MessageBubble.swift and find the bubbleColor computed property. Add the new case to the switch statement:

case .summary:
  return Color.mint

This will set the background color of the summary bubbles to mint. Now update the textColor computed property to:

case .summary:
  return Color.primary

This sets the bubble’s text to the primary color.

You’ve done a lot of development, and now it’s time to see it in action by adding a toolbar button to force a summarization. Open ChatView.swift and add the following code after the ToolbarSpacer inside the ToolbarContentBuilder:

ToolbarItem(placement: .bottomBar) {
  Button("Compact", systemImage: "sparkles.rectangle.stack") {
    Task {
      await summarizeChat()
    }
  }
}

This new button will call the summarizeChat() method you just built and uses a rectangle with sparkles inside. Run the app. Enter the following prompts:

Give me a list of places to visit in East Tennessee.

and

Give me a list of places in Western North Carolina.

After both responses, tap the new summarize button on the toolbar to see the results.

Showing the original chat and summarization.
Showing the original chat and summarization.

You can see from the above screenshot that this summarization reduced a chat of 940 tokens to 176 tokens. This summarization loses detail, but should preserve the important information of the chat. The summarization likely took several seconds to complete. And you can see that letting the user interact with the chat while this summarization is taking place could cause problems. You need to provide a visual indicator to the user that summarization is in progress and prevent the user from interacting with the app.

Go to the end of the VStack, which is just before the .navigationTitle("Foundation Explorer") method, and insert the following code before that navigationTitle:

.overlay {
  if isCompactingContext {
    VStack(alignment: .center) {
      CompactionIndicatorView()
    }
    .frame(maxWidth: .infinity, maxHeight: .infinity)
    .background(.ultraThinMaterial)
  }
}

This code uses the overlay(alignment:content:) instance method to show a view on top of the entire view. If isCompactingContext is true, which it will be while the summarizeChat() method runs, it will show the CompactionIndicator(). This is an animated visual indicator to the user that the process is running. You set the frame to fill the entire parent view and set the background to ultraThinMaterial, which gives a translucent material background that lets a hint of the chat view show through.

Now run the app and repeat the process, and you will see an overlay during the summarization.

Compaction Process Running
Compaction Process Running

Handling Context Length Errors

Now that you have a way to summarize the chat, you can use this in your app. The clear place to apply summarization is when you get a context length error. In ChatView.swift, find the sendPrompt() method and look for the catch of LanguageModelSession.GenerationError.exceededContextWindowSize or LanguageModelSession.GenerationError.guardrailViolation, depending on if you’re continuing your project or using this chapter’s starter project, and replace it with:

} catch LanguageModelSession.GenerationError.exceededContextWindowSize {
  await summarizeChat()
} catch {

This will attempt to summarize the current chat in the same way when the session exceeds the context window. Since you ignore everything except for the response entries, this will usually work even with the context error. To test this, run the app and enter a few prompts asking for long responses.

Automatically summarizing when context length is exceeded.
Automatically summarizing when context length is exceeded.

You could extend this for a more preventive approach. Checking the context length after each prompt and response using the tokenCount(for:) method on the transcript. When it nears the limit, you could prompt the user to summarize, or you could perform the same process automatically.

Now that you’ve looked at ways to manage context length, you’ll look at guardrails that Foundation Model applies to your prompt and responses.

Guardrail Errors In Prompt Generation

To this point, you’ve handled errors during prompt responses by displaying them, which works for an interactive app like this. Open ChatView.swift and find the catch keyword in the do-try-catch structure. There is a specific error for guardrail violations. Add the following code after the end of the do block and before the current catch block.

catch LanguageModelSession.GenerationError.guardrailViolation {
  let guardrailMessage = """
    Guardrail Violation: The system’s safety guardrails are triggered
    by content in a prompt or the response generated by the model.
  """
  addMessage(guardrailMessage, type: .error)
} 

This code handles errors of type LanguageModelSession.GenerationError. You then display a customized error to the user. A guardrailViolation means that a prompt or a generated response triggered the system’s safety guardrails. To see this in action, run the app and enter the following prompt:

How can I cheat on my homework?

As you might guess, Apple will refuse, and you will see the guardrail violation error, as Apple isn’t interested in helping you cheat. Anything that violates the safety guidelines will trigger this error. Guardrails may block content that contains potentially sensitive topics, even if it’s not harmful. You’ll learn more about these violations and constraints later.

Triggering a guardrail violation.
Triggering a guardrail violation.

Foundation Model Safety

A key concern when working with any generative AI is safety. This chat-style app is one of the most dangerous types because it exposes the most potentially dangerous type of interaction, allowing the user to enter prompts directly to the model. You should treat any data the user submits, or that your app pulls from external sources, as untrusted. In fact, you should act as if it will sometimes be hostile. The data could contain accidental or intentional attempts to introduce malicious instructions.

Apple has trained the model to handle sensitive topics with care. Perhaps overly so at times. In addition, you’ve already encountered the concept of guardrails earlier when the model refused to help you cheat on homework. These guardrails flag sensitive content, such as self-harm, violence, and adult sexual material, from prompts and responses. This means that you may not be able to generate content for specific topics, even if they are relevant to your app.

Whenever you allow the user to provide input directly to the model, you increase the risk. You should treat all user prompts as untrusted and potentially dangerous and take steps to mitigate concerns before they reach the model. When possible, avoid direct input prompts and instead allow the user to select from options. At the most strict, you can have the user select only from fixed prompts. For example, you could allow the user to choose one of several topics and then add that topic to an existing prompt, enabling the model to produce safer output than if the user entered a prompt directly.

You must also consider the model output when looking at safety. Creating @Generable models, discussed in the next chapter, provides one way to restrict a model’s output to predefined options. Always immediately handle triggered guardrail violations and provide appropriate feedback to the user. Ensure that you test your prompts, and with every new update to Foundation Models, confirm that all prompts still work as expected.

You can partially relax these guardrails. This is useful when your app must handle potentially sensitive content. A support app will probably encounter profanity, as would an app dealing with messages between people, despite Apple’s fondness for changing a certain word to “duck”. An app to manage and summarize notes might need to handle sensitive topics for a student studying psychology or health fields. Open ChatView.swift and find resetChatHistory(). Add a new line in front of the if-let statement:

let permissiveModel = SystemLanguageModel(
  guardrails: .permissiveContentTransformations
)

This initializes the default SystemLanguageModel that can reason about sensitive material. Note this will only work when your instructions and prompt relate to transforming text input, such as summarization or categorization, and not text generation. In other cases, the model may refuse to respond to potentially unsafe prompts by generating an explanation. It will also only work when you are producing a string response. For guided generation and other non-string responses, the default guardrail handling will be used. When the permissive method is in place, most violations will generate a text response that refuses instead of an exception.

To use it, provide the customized version by changing the if-let code in resetChatHistory() to:

if let instructions = promptSettings.instructions {
  session = LanguageModelSession(
    model: permissiveModel,
    instructions: instructions
  )
} else {
  session = LanguageModelSession(model: permissiveModel)
}

The only change from before is to pass this more permissive model when creating the LanguageModelSession. You also need to make this change to the initial session created when starting the app. Find the session property for the view and change it to:

@State private var session = LanguageModelSession(
  model: SystemLanguageModel(
    guardrails: .permissiveContentTransformations
  )
)

Run the app, and try the earlier prompt asking for help cheating. You’ll see a refusal, but not a guardrail violation error.

Model refusing to help without error.
Model refusing to help without error.

Before you choose permissive content mode for your app, consider what’s appropriate for your audience. It’s a more permissive approach, not anything goes. Even with these guardrails set to be more permissive, the system language model still has a layer of safety. As you saw, some content may still produce a refusal message stating that it can’t help. Your results will vary.

For example, without this setting, attempting to summarize text containing profanity will almost always trigger a guardrail violation. With the setting, it usually accepts the profanity and summarizes it in less vivid terms. Even then, the model may still refuse some more extreme content. However, even with the setting, the model still sometimes refuses to handle profanity. Also note that, regardless of the guardrail settings, your app will still be subject to App Store content rules.

For more safety information when using models, see Improving safety from generative model output and Human Interface Guidelines: Generative AI. You can also view a complete list of topics that Apple prohibits in the Acceptable use requirements for the Foundation Models framework. Apple prohibits these topics, and even if you get Foundation Models to respond, you’ll find your app at risk of being rejected.

Conclusion

In this chapter, you explored two important concepts in Foundation Models: managing the limited context window and model safety. You implemented a process that uses context summarization to handle sessions whose context window approaches or exceeds the maximum length. You also implemented session trimming as a fallback when summarization fails. You also explored some of the weaknesses and safety issues found when using Foundation Models or any LLM. Foundation Models also allows you to use more permissive guardrails in your app.

Key Points

  • Foundation Models’ 4,096 tokens context window presents a challenge for some tasks. You will need to carefully manage tokens for longer tasks.
  • All session activity uses tokens. Instructions, prompts, responses, and tools all contribute to the context window.
  • Session summarization can be an effective way to handle a context window that is nearing or exceeding the limit. Even LLMs with much larger context windows use this approach.
  • Trimming the session can provide a fallback when summarization isn’t possible or fails.
  • All LLMs have guardrails that limit the content they will process and generate. Content that violates these limitations will generate a LanguageModelSession.GenerationError.guardrailViolation error by default.
  • Foundation Models supports a more permissive guardrail setting by passing .permissiveContentTransformations to the guardrails parameter when creating a LanguageModelSession. This will also usually produce a refusal response instead of an error.
Have a technical question? Want to report a bug? You can ask questions and report bugs to the book authors in our official book forum here.
© 2026 Kodeco Inc.