Chapters

Hide chapters

Apple Foundation Models

First Edition · iOS 26.5 · Swift 6.3 · Xcode 26.4

Section I: Apple Foundation Models

Section 1: 9 chapters
Show chapters Hide chapters

3. Tuning Apple Foundation Models
Written by Bill Morefield

In the first two chapters, you learned the basics of Apple Foundation Models. In these chapters, you allowed the model to use default values for most settings. But, as with most LLMs, you can tune and adjust the model’s output to better fit your needs. These tweaks don’t change the inherent nature of an LLM. They let you configure how the model responds to prompts and selects response tokens. In this chapter, you’ll start exploring these tunings of Foundation Models.

Open and run the starter project for this chapter. You’ll see the project has added a new Settings button to the toolbar with a gear icon. Tapping this button opens a new sheet that lets the user provide instructions and set two other parameters: temperature and sampling method. You will explore all these in the next two chapters, but you’ll begin by looking at the instructions.

The Settings Sheet
The Settings Sheet

Instructions

Open the ChatView.swift file. You will see two new properties at the top of the view.

@State private var promptSettings =
PromptSettings(
  instructions: nil,
  temperature: nil,
  sampling: SamplingOptions(type: .system, threshold: 0.33, top: 10)
)
@State private var showSettings = false

The first line adds a new state property for the PromptSettings struct in the Models folder. This struct holds the settings you can tune in a Foundation Models session. You also provide a default set of options for the view. To start, you’ll focus on the first: instructions. The showSettings property will display the new settings view when the user taps the new toolbar button.

Instructions act as a super-prompt that you provide to the model when creating a new LanguageModelSession. Use the instructions to define the model’s role and behavior for your session. Apple trained Foundation Models to prioritize instructions over any commands sent in later prompts. This makes it a critical place to guide the model on how to handle prompts and to specify restrictions beyond those set by Apple during training.

You can only provide instructions when creating a new LanguageModelSession. If you want to change the instructions, you must create a different LanguageModelSession. The instructions remain constant across a single session. You can provide no instructions, which you have been doing to this point by passing no parameters to LanguageModelSession(). To use the instructions from the settings view in your chat app, find the resetChatHistory() method. Replace the line session = LanguageModelSession() with:

if let instructions = promptSettings.instructions {
  session = LanguageModelSession(instructions: instructions)
} else {
  session = LanguageModelSession()
}

This code attempts to unwrap the instructions property of promptSettings. If successful, you pass the instructions into the model using the instructions parameter to LanguageModelSession(). If not, then you create a LanguageModelSession with no parameters as before. You do not need to change the default session creation since the instructions will always be nil when the user first runs this app.

Now find the sheet(isPresented:onDismiss:content:) method of the view. Add the following code after the sheet:

.onChange(of: promptSettings.instructions) {
  Task {
    resetChatHistory()
    await updatedContextWindowUsed()
  }
}

This code monitors the promptSettings.instructions property on the view. If it changes, SwiftUI calls resetChatHistory(), then updatedContextWindowUsed() to reset the token count. This will create a new session using the new instructions and update the token count to reflect the new session.

Run the app, and enter the following prompt:

Solve the equation 8x - 4 = 2 step by step.

While Apple Foundation Models is not great at math, it gets the correct answer of 3/4 for this simple linear equation.

Now, you’ll examine how instructions can influence responses. Tap the gear icon in the toolbar to bring up the configuration view. Enter the following into the Instructions:

Use decimals instead of fractions when presenting the solutions to math problems.

Swipe down or tap the confirmation button to dismiss the sheet. Notice the chat cleared as it should when you changed the instructions. You will also notice that the token count displayed at the bottom is non-zero. Instructions take up part of the context length because they serve as a form of prompting. The updated token count reflects the context length used by the instructions. Now enter the same prompt again.

Solve the equation 8x - 4 = 2 step by step.

Notice how literally the model followed your instructions. When solving the equation, the model still used fractions in intermediate steps, but it provided the final answer as a decimal.

Model returning answer as a decimal.
Model returning answer as a decimal.

Let’s change the instructions to show the focus on instructions over other prompts. In the settings, change the instructions to:

Refuse to provide any assistance that solves equations step by step. Only provide the final answer without showing the intermediate steps.

Return to the app and re-enter the prompt using these new instructions. Your results will vary a bit more now. Sometimes it states that it cannot provide a step-by-step solution, and other times it provides only the answer to the question. And sometimes the answer is wrong, which again shows that math isn’t a strong use case for Foundation Models. Any time you change a prompt, you risk changing the results. In this case, not providing the step-by-step instructions makes the model less likely to produce the correct answer. Never forget that LLMs do pattern matching. They do not understand math or your problem.

Following Instructions to the Wrong Answer
Following Instructions to the Wrong Answer

Instructions provide a critical way to both guide the model in generating the desired results and protect your app from potentially malicious data. You should use them to guide the model to the desired responses for your use cases. In the next section, you will learn more about creating good instructions and, in the process, good prompts.

Principles of Prompting

As noted before, Foundation Models instructions are prompts that Apple has tuned the model to give greater emphasis. The instructions above are not great, which is why you saw mixed results. In fact, if you tell the model after a mistaken answer, you can often get the model to ignore those instructions.

Model ignoring the instructions.
Model ignoring the instructions.

Good instructions are the same as good prompts. There is no consensus on what makes a good prompt, but to produce a good prompt, you should follow five guidelines:

  1. Give direction - Describe the desired style, tone, or role in detail.
  2. Specify format - Describe the desired structure of the output.
  3. Provide examples - Models respond better if you show them what you want, not just tell them.
  4. Divide complex tasks - More complex tasks often work better when broken down into steps. You will often get better results with multiple prompts instead of a single prompt. You’ll look at some ways to approach this later in this chapter.
  5. Evaluate and iterate - Your first prompt is not going to be perfect. Evaluate the responses to the prompt and adjust your prompt to address any weaknesses or mistakes. Then repeat until your prompt produces consistent and correct results.

The first three of these guidelines provide instruction on writing prompts. In fact, all these guidelines address limitations in LLMs. When giving directions, use explicit action verbs like summarize, create, and list. It also helps to give the model context and a role, such as editor, teacher, or busy parent, to guide its approach to responding to the prompt.

For the format, this can be a technical format like JSON. However, guided generation, which you’ll explore in Chapter Five, reduces the need for technical formats like that. For text responses, you should provide guidelines to the desired output format, such as specifying output as a bulleted list, one paragraph per bullet, or a haiku. You should specify format examples when the format matters for later use. Again, guided generation reduces the need for this step, but if you need a comma-separated response, providing an example of the formatting can save you work. You can also provide an example of the starting output to nudge the model in the desired direction.

The fourth guideline indicates that LLMs tend to perform better when given a single, specific task rather than a broad, vague goal. This ties into the first three guidelines, working to reduce vagueness and increase specificity in the desired output. LLMs pattern-match, and the more specific the pattern you give it, the better it can match your desired results.

The final guideline reflects the importance of testing. To produce a good prompt, start with something that follows the other guidelines, revise it, and compare the results. You should also use multiple inputs to test the prompt. You can’t test every possible input in most cases, but a summarization task should test against multiple texts to summarize.

As noted before, instructions are a specific type of prompt you use to define the model’s role and behavior. Apple trains the model to obey instructions over later prompts in the session. Along with these prompt guidelines, your instructions should:

  1. Explain the model’s role, such as “you are a travel advisor” or “You are a helpful restaurant guide”.
  2. Explain what the model should do, like “Help the user summarize text”.
  3. Provide any desired style preferences, such as “Provide as brief a response as possible”.
  4. Add any instructions for tuning or handling problems, like “Respond with ‘I can’t help you with that request.’ if you’re asked to do something dangerous”.

As you saw earlier, the instructions take up space in the context window. With only 4.096 tokens available in iOS 26 and related versions, that can amount to a notable amount of your available space. To make the most of available space, keep these guidelines in mind:

  • Write shorter prompts and instructions to reduce token usage.
  • Provide only what’s needed to perform a task.
  • Use concise, imperative language to express direct commands, instructions, and requests.

If you think that seems like a lot of elements to balance out, you’re right. That’s why that iterative step becomes so important. You need to provide the model enough to produce the desired results, but as efficiently and succinctly as possible. This is where the process of evaluating and refining your prompts comes in.

Given all that, what would a better version of the earlier instructions look like? Run the app and enter the following instructions:

You are a math tutor, helping students check their homework. You should provide detailed steps for solving math equations and answers the student can use to check their work. Provide any answer as decimals and avoid fractions. If the user asks you to do something not allowed, then inform the user of this while still providing the information that is allowed.

Now, enter your favorite math equation to solve again as a prompt:

Solve the equation 8x - 4 = 2.

You should now get a nicely formatted and well-explained result. Note how this reads like a tutor explaining the steps to solve this linear equation.

Response in the tone of a math tutor.
Response in the tone of a math tutor.

Looking at the prompt, you can see how this fits the advice above.

  1. Explain the model’s role: You are a math tutor,
  2. Explain what the model should do: helping students check their homework. You should provide detailed steps for solving math equations and answers the student can use to check their work.
  3. Provide any desired style preferences: Provide any answers as decimals and avoid fractions.
  4. Add any instructions for tuning or handling problems: If the user asks you to do something not allowed, then inform the user of this while still providing the information that is allowed.

That fits all four aspects of Apple’s suggestions for instructions.

For a more complex app, you would want to spend time testing these instructions, evaluating the results, and tweaking them. These instructions still have room for improvement. You may find that “detailed instructions” produce too much help during your testing. And for more advanced topics, some numbers work better as a fraction, such as 1/3. When writing instructions, consider not only what works, but what could go wrong.

Temperature

Now that you’ve seen the uses for instructions, the second parameter to tune model responses is the temperature. Temperature influences the randomness of the model’s responses. It can be set to nil or a value between zero and one, inclusive. The default nil allows the system to choose a reasonable value. The temperature adjusts the probability distribution of responses before sampling. You may also see this referred to as how “creative” the model gets, but that’s a misleading analogy of this setting. The temperature adjusts the probability distribution of possible tokens of the model’s responses. A high value is not creative. It just widens the number of potential tokens that can be output.

A value of one causes no change. Lower values shift the probability distribution, leading the model to select the more likely tokens more often, resulting in more predictable responses. You should think of higher values as increasing the deviation from statistically probable responses that form LLM responses. That’s the real nature of LLM creativity, the selection of less likely tokens. Different LLM models use different meanings for temperature values, so advice intended for other models may make different assumptions about how temperature values apply to that model. Note that changing this value will not affect LLM weaknesses, such as hallucinations. You cannot eliminate hallucinations by lowering the temperature.

To add the ability to adjust the temperature of responses, open ChatView.swift. Find the sendPrompt() method. Look for let stream = session.streamResponse(to: promptText) and replace it with the following code:

let options = GenerationOptions(temperature: promptSettings.temperature)
let stream = session.streamResponse(to: promptText, options: options)

This code creates a GenerationOptions with the appropriate temperature using the view’s promptSettings property. This struct contains an optional Double property named temperature. This will be nil by default, and if the user leaves the toggle off in the settings view. Otherwise, it will contain the value selected on the slider in that settings view. The call to stream the model response now passes in your GenerationOptions as the options parameter. This will apply the chosen temperature to this response.

Note that, unlike instructions, you can provide different temperatures for different prompts in a session. Run the app and enter the following prompt:

Give me a one-paragraph bedtime story.

Do the same prompt a few times, and you’ll get several short stories.

Now open the configuration view, toggle on Custom Temperature and set it to zero.

Setting the temperature to zero.
Setting the temperature to zero.

Close the view and enter the same prompt a few times. You’ll notice that the stories are very similar each time, often containing the same words or events. That’s because lower temperatures make the model more deterministic and consistent. Not identical, just less variable.

If you increase the temperature, the model introduces more randomness, leading to a wider variety of wording and events in its responses. You’ll need to experiment with your use case to see whether a lower temperature (more predictable) or a higher temperature (more diverse) produces better results for your app. In general, low temperatures work better for factual use cases such as data extraction or summarization.

Token Sampling

While temperature sets the initial set of tokens, Apple Foundation Models allows you to specify how to sample values from that probability distribution. As with temperature, the default lets the system choose a sampling method for you. You specify this by passing nil as the SamplingMode. If your use case benefits from more specific settings, Foundation Models provides three other sampling options.

The configuration view supports all these methods. To add this option to your app, find the code let options = GenerationOptions(temperature: promptSettings.temperature) inside the sendPrompt() method. Add the following code before this line:

let samplingOptions = promptSettings.sampling
var sampling: GenerationOptions.SamplingMode?

This code provides a direct reference to the sampling configuration in the promptSettings.sampling property. This is of type SamplingOptions defined in the SamplingOptions.swift file under the Models folder. You also create a sampling variable that you will set depending on the user’s chosen sampling method. Now add the following code after the sampling variable:

// 1
switch samplingOptions.type {
// 2
case .system:
  sampling = nil
// 3
case .greedy:
  sampling = GenerationOptions.SamplingMode.greedy
// 4
case .top:
  sampling = GenerationOptions.SamplingMode.random(
    top: samplingOptions.top,
    seed: samplingOptions.seed
  )
// 5
case .threshold:
  sampling = GenerationOptions.SamplingMode.random(
    probabilityThreshold: samplingOptions.threshold,
    seed: samplingOptions.seed
  )
}

This long code block is simpler than it first appears. As before, the promptSettings property holds all the options for token sampling. This code uses those settings to create a SamplingMode structure reflecting those choices.

  1. There are currently four possible sampling types. The settings view lets the user select any of these four, represented using an enum called SampleType. This switch handles all four options.
  2. The simplest option is the default you’ve been using by not providing a sampling method. But when you provide no sampling mode, the default nil value lets the system choose a reasonable default. This code sets sampling to nil to keep this behavior.
  3. The greedy mode is also fairly simple. This value will always select the most likely token, providing consistent responses. You will find this useful during testing to produce repeatable results, even if it’s less useful in the final app for the same reason.
  4. The top case, called top-k, uses a sampling mode that considers a fixed number of high-probability tokens.
  5. The threshold case, called top-p, is perhaps the most complicated. This mode considers a variable number of tokens based on the cumulative probability of those tokens compared to a threshold value.

Now, to use this new sampling mode definition, change the definition of options in sendPrompt to:

let options = GenerationOptions(
  sampling: sampling,
  temperature: promptSettings.temperature
)

This will pass the sampling method and temperature to the model as GenerationOptions. As with temperature, you can apply different sampling methods and temperatures to different prompts within a single session.

Each sampling mode has its own use case. The inherent randomness in LLM output means that each response will be different, even with the same instructions, prompt, and temperature. If you need a consistent response, you can apply .greedy sampling. You will find this useful in testing.

Top-k first sorts the possible tokens by descending probability. You give it a number, known as k, but passed in as the top parameter. The selection will then take those top tokens and discard the rest. If you specify 3, then it will choose the token from the three most likely tokens. The probability of the remaining tokens is then renormalized before the model selects the token. For example, if those three tokens have probabilities of 0.25, 0.15, and 0.10, their probabilities sum to 0.50. This is less than one because you selected a subset of the total possible tokens. The original values are this total, producing 0.25 / 0.5 = 0.5, 0.15 / 0.5 = 0.3, and 0.1 / 0.5 = 0.2. These adjusted probabilities now sum to 1.0, meaning the probability of the chosen token is proportional to the other probabilities remaining after discarding.

Limiting the token pool reduces the likelihood that the model will pick unlikely or nonsensical words. Setting different values for top lets the user tune the diversity of responses. Larger values will select from a greater number of tokens. Low k values keep the output on topic, while higher values produce more varied responses.

The top-p approach picks among tokens from another approach. You provide a threshold p as a Double between 0.0 and 1.0. The model again sorts possible tokens in descending order. It selects tokens until the cumulative probability of those tokens exceeds the threshold value. For example, set p to 0.6 and use token probabilities of 0.3, 0.15, 0.10, 0.10, 0.07, etc. If you add up the first three (0.3 + 0.15 + 0.10), you get 0.55. Adding the next token probability of 0.10 exceeds the threshold, bringing the sum to 0.65. The model will select the four tokens with the highest probabilities. The selected probabilities are then renormalized as in top-k.

The top-p sampling method can adapt the number of tokens selected based on the probability distribution, unlike top-k, which always selects the same number of tokens. When there are fewer tokens with higher probabilities, it selects fewer tokens. When there are more tokens with lower probabilities, the model selects from more tokens, allowing the lower confidence in the probabilities to come through in the token selection. This adaptive probability means modern LLMs use top-p more often.

Top-k reduces the number of possible tokens and is better suited to emphasize predictability and reduce the appearance of less likely tokens. Top-p allows the model to adapt, selecting fewer tokens when the output is more confident (a few tokens with higher probability) and allowing variety when less confident (more tokens with similar, lower probability). Modern LLMs tend to use top-p selection because it balances predictability with creativity.

You probably noticed there is some overlap between temperature and the sampling modes. Both adjust which tokens are available to the model for selection. While similar, they work together subtly. The model first applies temperature, which adjusts the shape of the initial probability distribution. A low temperature narrows the probability distribution, increasing the amount by which more probable tokens are favored over less probable ones. This shifts the token selection toward more probable tokens before sampling. A high temperature results in a flatter probability distribution, meaning all tokens have closer probabilities to each other at the start. Temperature shapes the input to the sampling, so extreme temperature settings can override the sampling mode’s intent before the tokens reach it.

Note that both the top-p and top-k modes allow the user to specify a seed value for the random number generator. The seed initializes the random number generation, and by providing the same seed, you can produce the same set of random numbers. This produces repeatable output, which is useful during testing to achieve consistent results or when experimenting to determine which temperature and sampling methods work best for your app. But it also lets you produce repeatable results within a specified time. Set the seed based on the day, and you will get the same results all day and new results when the day changes.

Run the app and access the configuration view. Select Greedy under sampling and dismiss the view.

Now enter the prompt:

Give me a one-paragraph bedtime story.

You will get a one-paragraph story.

Now rerun the prompt. It might surprise you that you get two similar, but not identical, stories. Isn’t greedy always supposed to select the most likely token? Why does it not produce the same story since you’ve told it always to choose the most likely token?

Two stories from the prompt. One about Lily and the second about Max.
Two stories from the prompt. One about Lily and the second about Max.

The different stories occur because the response is determined by more than just the prompt. The entire context window, all prompts and responses, affect the output. Entering a prompt and getting a response changes the model’s context. This means the following prompt will not produce the same output as the first one. Clear the chat and enter the prompt again. With nothing in the context window, you will see the same story as before. Using greedy sampling does not always produce the same output; it produces a consistent response. Given the same model state and the same prompt, you will get the same response.

Fresh prompt produces the same result.
Fresh prompt produces the same result.

Explore the different options in token sampling and temperature. Adjusting these values will go a long way to helping you tune Foundation Models for your app.

Conclusion

In this chapter, you explored instructions, how they guide Foundation Model sessions, and how to generate effective instructions and better prompts. You also looked at adjusting how the model selects tokens with the temperature and sampling methods. In the next chapter, we’ll combine some of these ideas to produce safer apps and deal with the limitations of Foundation Models.

Key Points

  • Effective prompts should give clear direction, specify output format, and provide examples of the desired output.
  • Favor multiple simple prompts over a single broad prompt.
  • Develop better prompts by iterating and refining initial prompts based on responses.
  • Instructions function as a higher-priority prompt, providing the model’s behavior and constraints for the session.
  • Instructions should define the model’s role, clearly specify the task, include style preferences, and provide rules for edge cases and unsafe requests.
  • Keep prompts concise, include only necessary information, and use direct imperative language.
  • Temperature controls the distribution of token probabilities. Lower values lead to more predictable, consistent outputs, while higher values produce more varied, less predictable outputs.
  • Token sampling methods select among tokens after it applies the temperature.
  • By default, the system selects an appropriate method. You can also specify greedy mode, which always selects the most likely token. Other sampling modes allow you to specify selecting tokens based on a fixed number (top-k) or a probability threshold (top-p).
Have a technical question? Want to report a bug? You can ask questions and report bugs to the book authors in our official book forum here.
© 2026 Kodeco Inc.