Chapters

Hide chapters

Apple Foundation Models

First Edition · iOS 26.5 · Swift 6.3 · Xcode 26.4

Section I: Apple Foundation Models

Section 1: 9 chapters
Show chapters Hide chapters

7. Extending an App with Foundation Models
Written by Bill Morefield

In the earlier chapters of this book, you explored aspects of using Apple Foundation Models in isolation. You built apps purpose-designed to help explore and understand the framework’s features. To close out this book, you’ll look at a simple, real-world application that lets users record voice notes. You will then explore how using machine learning, in general, and Apple Foundation Models, in particular, can take this application from a useful but basic tool into a much more useful and powerful app.

A common mistake is treating adding machine learning and artificial intelligence as the goal, when it’s really a tool to improve your app’s user experience. Outside the challenge of developing good prompts to get the results you need, Foundation Models code will often be the simplest code in your app. Preparing the data to send to the model and then presenting the results back to the user in ways to provide insight and a better understanding are the real challenges.

The starting voice recording app.
The starting voice recording app.

Open the starter project for this chapter and run it on a device or in the simulator. You’ll see that the app allows the user to record short voice notes and provides an interface to review existing notes and delete them. The app also includes a few notes as seed data that you can use or delete. In the top-trailing edge of the app, you will see a button you can use when running through Xcode that allows you to restore these sample notes if you delete them.

Run the app and record a new note. You will have to allow access to the microphone to record. The appropriate permissions and values for this prompt are already in the app.

Allow microphone permissions.
Allow microphone permissions.

While this is a full-featured app, it provides no value beyond recording and playback. The information in the recordings remains locked into the audio. To make a more valuable app, you can give users the ability to examine and surface information from those recordings without having to listen to every note. By now, you might already be thinking of ways that Foundation Models could do this. While there are LLMs that can work directly with audio data, Foundation Models only works on text. But as the 26.0 versions of the operating systems introduced Foundation Models, it also introduced a service called SpeechTranscriber, perfect for this task.

Transcribing Audio into Text

Transcribing is converting audio into text. This task has been traditionally done by people using shorthand or other abbreviated writing methods to record speech at high speed. With advances in machine learning, computers can produce high-quality transcriptions of text on a local device. Apple introduced SpeechTranscriber in the same versions that brought Foundation Models to local devices. SpeechTranscriber is a speech-to-text transcription module built for normal conversations and transcription.

There is a limitation in SpeechTranscriber. It only works on actual devices, not in the simulator.

SpeechTranscription not available inside the simulator.
SpeechTranscription not available inside the simulator.

If you do not have a device to run the app on for this chapter, you can take advantage of the Designed for iPad option to run the iPad version of the app on your Mac, which does provide SpeechTranscriber support. Do this by selecting the My Mac (Designed for iPad) option as the device to run the app on.

Running the app on a Mac.
Running the app on a Mac.

To run an app through Xcode on your Mac, you will need to go to the VoiceNotes target and assign a Team under the Signing & Capabilities tab. A free account should work for this chapter. You may also need to add a unique bundle identifier.

Create a new Swift file under the empty Services folder named SpeechTranscriptionService.swift. This file will contain the service to perform transcription of the voice recording once the recording completes. Replace the contents of the file with:

import AVFoundation
import Foundation
import Speech

enum SpeechTranscriptionError: LocalizedError {
  case unavailable
  case authorizationDenied
  case unsupportedLocale
  case emptyResult

  var errorDescription: String? {
    switch self {
    case .unavailable:
      "SpeechTranscriber is not available on this device."
    case .authorizationDenied:
      "Speech recognition access is needed to transcribe voice notes."
    case .unsupportedLocale:
      "SpeechTranscriber does not support the current language."
    case .emptyResult:
      "No speech was detected in this recording."
    }
  }
}

This code produces an error enum that you will use to provide feedback to the user if anything goes wrong during transcription. Now add the following after the SpeechTranscriptionError enum:

struct SpeechTranscriptionService {
  // 1
  func transcribeAudio(at url: URL) async throws -> String {
    // 2
    guard await requestSpeechAuthorization() else {
      throw SpeechTranscriptionError.authorizationDenied
    }
    
    // 3
    return try await transcribeWithSpeechTranscriber(at: url)
  }
}

Here’s how this code sets up transcription:

  1. The transcribeAudio(at:) method will take a URL to the audio file and return a string with the transcribed text. You mark the method as asynchronous and throws.
  2. As with many features under Apple OS’s, speech transcription is locked behind a user permissions request. This method will verify permission and prompt for permission if needed. If permissions are not given or previously denied, then it will throw the SpeechTranscriptionError.authorizationDenied error. You will implement this method in a moment.
  3. If the app has the needed permissions, it will then call a method transcribeWithSpeechTranscriber(at:) to perform the transcription, which you will also implement very soon.

With that setup, you need to implement the method to verify and prompt for permissions. Add a new method at the end of the SpeechTranscriptionService struct:

private func requestSpeechAuthorization() async -> Bool {
  await withCheckedContinuation { continuation in
    SFSpeechRecognizer.requestAuthorization { status in
      continuation.resume(returning: status == .authorized)
    }
  }
}

This code defines the requestSpeechAuthorization() method to manage permissions. The core is the SFSpeechRecognizer.requestAuthorization code, which you need before you attempt to perform any recognition tasks. Otherwise, it will fail. Despite using the more modern SpeechTranscriber framework, you still request permission through the older SFSpeechRecognizer framework. You should call this code on the main thread before you access speech recognition for the first time. The first time you call it, the system will prompt the user to grant or deny permission and then remember that choice. Make sure to select Accept in your app, or you will need to change it in the app’s settings. Here, we return true if the returned status is authorized. Otherwise, the method will return false.

You must add the NSSpeechRecognitionUsageDescription key to your Target or the app will crash when you attempt to use this method. To do this, go to the Project for the app in Xcode and select the VoiceNotes target. Go to the Info tab, and you will see the existing list of properties. Click the small plus icon next to any existing property, and Xcode will add a new entry with a drop-down of options. Scroll down and find Privacy - Speech Recognition Usage Description and set the value to:

Voice Notes needs speech recognition access to transcribe recordings.

This completes the steps to seek and receive the user’s permissions to access speech recognition and, therefore, transcribe audio. Go back to SpeechTranscriptionService.swift and add the following method to the end of the SpeechTranscriptionService struct:

private func transcribeWithSpeechTranscriber(at url: URL) async throws -> String {
  guard SpeechTranscriber.isAvailable else {
    throw SpeechTranscriptionError.unavailable
  }

  guard let locale =
    await SpeechTranscriber.supportedLocale(equivalentTo: .current) else {
    throw SpeechTranscriptionError.unsupportedLocale
  }
}

This code first ensures that SpeechTranscriber is available and supports the current locale. As noted earlier, this library requires an operating system version 26.0 or later. There are libraries for older operating systems, but since you need the same minimum for Foundation Models, support for older versions provides limited benefit to this app. If SpeechTranscriber is unavailable, the code throws the appropriate error. The locale check ensures that the library supports the device’s current language and stores that locale in the locale variable.

To finish the setup for transcription, continue the transcribeWithSpeechTranscriber(at:) method with the following code:

let transcriber = SpeechTranscriber(locale: locale, preset: .transcription)
let modules: [any SpeechModule] = [transcriber]
try await prepareAssets(for: modules)

You first create an instance of the SpeechTranscriber class, passing in the locale saved earlier and a preset indicating you want the configuration for basic, accurate transcription. Then you define an array containing this instance and pass it to another method that prepares the assets needed for speech transcription. You will implement that method soon. Continue the method with:

// 1
let audioFile = try AVAudioFile(forReading: url)
// 2
let resultsTask = Task {
  // 3
  var finalText = ""

  // 4
  for try await result in transcriber.results {
    // 5
    guard result.isFinal else { continue }
    // 6
    let text = String(result.text.characters)
      .trimmingCharacters(in: .whitespacesAndNewlines)
    // 7
    guard !text.isEmpty else { continue }

    // 8
    finalText += " " + text
  }

  // 9
  return finalText
}

This code handles the core of the transcription process:

  1. You begin by loading the audio file at the URL passed to the transcribeWithSpeechTranscriber(at:) method.
  2. The remainder of this code block sets up a Task to run the transcription. Note that this will not run immediately. You will run it later in the method.
  3. You create an empty string to hold the transcription results
  4. You should recognize the pattern of processing an asynchronous stream from managing streamed responses from Foundation Models earlier in this book. This code handles the asynchronous stream of transcription results within the for loop.
  5. The result.isFinal flag indicates that the transcription has finalized the text. The continue statement will move to the next loop iteration of the stream.
  6. The result.text property contains the transcribed segment. You trim any whitespace and new line characters and set the results to the text variable.
  7. This guard statement will continue to the next segment of the stream if the text is empty.
  8. You append the returned text to the current transcript. You append a space before adding text to separate each block from the others, as the blocks tend to fall on individual sentences and rarely in the middle of words.
  9. Once the stream completes, you return the accumulated finalText from the Task to the caller.

With this task set up to handle the stream, you can finish the method.

// 1
let analyzer = SpeechAnalyzer(modules: modules)
// 2
try await analyzer.start(inputAudioFile: audioFile, finishAfterFile: true)
// 3
let transcript = try await resultsTask.value
  .trimmingCharacters(in: .whitespacesAndNewlines)

// 4
guard !transcript.isEmpty else {
  throw SpeechTranscriptionError.emptyResult
}

return transcript
  1. You create an instance of the SpeechAnalyzer class, passing the single module you’ve prepared to it.
  2. You call the asynchronous start(inputAudioFile:finishAfterFile:) method, passing in the audio file to be transcribed along with a flag telling SpeechAnalyzer that the work is finished after the audio file has been fully processed. If you had set finishAfterFile to false, the stream would remain open, which can be useful for live transcription while a recording is in progress.
  3. This step performs the actual transcription using the resultsTask Task you set up earlier. It takes the value returned from the task and sets that result as the transcript.
  4. If the transcript is empty, then either something went wrong in the audio, or the transcription couldn’t find usable text in the audio to transcribe. Either way, you return an error about the empty result. If the transcript is non-empty, the method returns its text.

The last step to implement transcription is the prepareAssets(for:) method. Part of using SpeechAnalyzer is ensuring the right assets are available through the AssetInventory class. This class manages the necessary assets for transcription or other analyses. Add the following new method to the end of SpeechTranscriptionService:

private func prepareAssets(for modules: [any SpeechModule]) async throws {
  switch await AssetInventory.status(forModules: modules) {
  // 1
  case .installed:
    return
  // 2
  case .downloading:
    guard let request =
      try await AssetInventory.assetInstallationRequest(
        supporting: modules
      ) else {
      throw SpeechTranscriptionError.unavailable
    }
    try await request.downloadAndInstall()
  case .supported:
    guard let request =
      try await AssetInventory.assetInstallationRequest(
        supporting: modules
      ) else {
      throw SpeechTranscriptionError.unavailable
    }
    try await request.downloadAndInstall()
  // 3
  case .unsupported:
    throw SpeechTranscriptionError.unavailable
  // 4
  @unknown default:
    throw SpeechTranscriptionError.unavailable
  }
}

In this app, recall that you call SpeechTranscriber with the current locale and target it for general transcription. Doing so requires machine-learning models downloaded from Apple’s servers and managed by the system. The system handles downloading these models. Once downloaded, these models are available for all other apps and automatically updated. The method begins by getting the status property of AssetInventory for the modules passed into the method. For multiple modules, the status will return the least ready module’s status. The case statement checks the response in descending order of readiness for use. For these cases:

  1. If the assets are installed, everything is ready, and you return.
  2. For the downloading and supported case, you initiate an assetInstallationRequest(supporting:) passing in the requested models. This will return an installation request object, which you use to initiate the asset download and monitor its progress. If the return is nil, you throw an unavailable error. Otherwise, you call the downloadAndInstall() method on the returned installation request object to download and install any assets not already on the device. This method will return when the request either succeeds or fails.
  3. For the unsupported case, you again throw the unavailable error. This is the state the app returns when run in the simulator, as it does not support transcription.
  4. You use the @unknown default case for open system enums, which are often brought in from older frameworks. There is a chance that Apple may add new cases in the future, and by using this, we can gracefully handle any future operating system inclusions, along with receiving a compiler warning when updating the code.

This completes the implementation of the transcription service. With that part in place, you can now update the app to use this to transcribe the user’s recordings in the next section.

Implementing Note Transcription

Open VoiceNoteStore.swift under Models and add a new property after the player property near the top of the class:

private let transcriptionService = SpeechTranscriptionService()

This creates an instance of your SpeechTranscriptionService that the note store can use. Find resetSampleNotesForTesting(), which is surrounded by a #if DEBUG conditional, and add the following new method after the #endif line:

func transcribeRecording(_ note: VoiceNote) async -> String? {
  // 1
  guard note.transcript?.isEmpty != false else { return note.transcript }
  guard !transcribingNoteIDs.contains(note.id) else { return nil }

  // 2
  transcribingNoteIDs.insert(note.id)
  defer {
    transcribingNoteIDs.remove(note.id)
  }

  do {
    // 3
    let transcript =
      try await transcriptionService.transcribeAudio(at: url(for: note))
    updateTranscript(transcript, for: note.id)
    return transcript
  } catch {
    // 4
    permissionMessage = error.localizedDescription
    return nil
  }
}

Here’s what this code does:

  1. You first check to see if the note has an existing transcript, and if found, you return it. The transcribingNoteIDs on this class contains a Set of notes in the process of transcription. If this note’s id is in that set, it returns nil because the note is already in processing but not yet available. In practice, this case should not occur.
  2. You then add the note to the transcribingNoteIDs, showing its processing. The user interface also uses this property to show the user when note transcription is in progress. As you’ve seen, there are several steps, and the first transcription could take some time as assets download. You use the defer keyword to ensure the id gets removed from the set when the method completes in any way.
  3. This calls the transcribeAudio(at:) method on the service you built through the transcriptionService instance you created. It then calls the updateTranscript(_:for:) method, which updates the note with the transcript returned by the service.
  4. If anything goes wrong, the app sets the permissionMessage property on the view to the error. Any time this property is set within the store, ContentView displays the message to the user.

Now, to wire up the transcription to occur after the recording finishes, find the stopRecording() method in the store. Add the following code to the end of the method:

Task {
  await transcribeRecording(note)
}

This will call the new transcribeRecording(_:) method to transcribe the note when the recording completes, and saves.

Finally, you can put this to the test. Build the app and run it either on a physical device or on your Mac with the My Mac (Designed for iPad) option. Recall that the transcription will not work in the simulators and you must run it on a device or a Mac. On the first run for a device, it will prompt you to allow permissions.

Request for permissions to use speech recognition.
Request for permissions to use speech recognition.

Record a short note and watch the transcription appear after a pause.

A voice recording turned into a transcript.
A voice recording turned into a transcript.

It would be useful to have a way to manually retry a transcription if something went wrong the first time. This will also help during development, as you might record on a device that didn’t support transcription, and for sample recordings that don’t contain anything other than the recordings.

Sample recordings have no transcript.
Sample recordings have no transcript.

To add this, open VoiceNoteTranscriptSection.swift. Within the VoiceNoteTranscriptSection view, find the final else condition inside the view containing a Text view reading No transcript is available for this recording yet.. After this text view, add the following code:

Button {
  Task {
    await store.transcribeRecording(note)
  }
} label: {
  Label("Transcribe", systemImage: "text.bubble")
}
.buttonStyle(.borderedProminent)

This code adds a new button to the details of a note under the Transcript section. This button calls the transcribeRecording(_:) method you created for this note, which will result in a transcript. To see it in action, run the app again and tap on any of the sample recordings. Tap on the Transcribe button, and soon the transcription will appear.

Transcribing on of the sample recordings.
Transcribing on of the sample recordings.

You’ll also add a toolbar button that lets the user create a transcript on demand. Open VoiceNoteDetailView.swift and find the toolbar modifier on the ScrollView. Add the following code to the top of the ToolbarItemGroup closure:

if note.transcript == nil {
  Button {
    Task {
      await store.transcribeRecording(note)
    }
  } label: {
    Image(systemName: "text.bubble")
      .accessibilityLabel("Produce Transcript")
  }
  .disabled(store.transcribingNoteIDs.contains(note.id))
}

If the note does not have a transcript, then this code will show a button with a text bubble icon. Tapping the button initiates generating the transcription. The button calls the transcribeRecording(_:) on the VoiceNoteStore for the view to perform the transcription. It also checks to see if the note is being transcribed using the store’s transcribingNoteIDs property to disable the button during the transcription process.

Now with toolbar button to create transcripts.
Now with toolbar button to create transcripts.

These created transcripts provide value on their own. While the transcription will not be perfect, it’s easier to glance at a transcription than scan through a recording when on the go. In the next section, you will begin examining this by adding search ability for these transcriptions to the app.

Searching Transcriptions

Open ContentView.swift. Add a new state property to the view after store to hold the search text:

@State private var searchText = ""

To add the search box and tie it to this new property, find the navigationDestination modifier on the view and the following code after it:

.searchable(text: $searchText)

This modifier automatically adds search capability to your app when applied to a navigation element, handling the visual and layout work. You pass a binding to the searchText property to link the text the user enters to that property. Now that you can allow the user to enter search data, add the following new computed property after searchText:

var visibleNotes: [VoiceNote] {
  // 1
  if searchText.isEmpty {
    return store.notes
  }
  
  // 2
  return store.notes.filter { note in
    // 3
    if let transcript = note.transcript {
      // 4
      transcript.localizedStandardContains(searchText)
    } else {
      false
    }
  }
}

This computed property will handle filtering when the user enters text in the search field:

  1. First check if the searchText property is empty. If so, then it returns the full list of notes.
  2. You will use the filter(_:) instance method on the store.notes collection. This takes a predicate inside the closure and returns a new array containing only the members of the original set for which the predicate returns true. Within the closure, you will access this element through the note variable.
  3. You first attempt to unwrap the transcript property of the note. If this fails because the note has no transcript, you ignore it for the search by returning false.
  4. When a transcript exists, you use the localizedStandardContains(_:) instance method on the transcript. This returns true if transcript contains the searchText. This method performs a locale-aware search that ignores case and accent differences, meaning it will catch reasonable variations of searchText within transcript.

One more change so the app will display the list from the computed property. Find the line ForEach(store.notes) { note in inside the view and replace it with:

ForEach(visibleNotes) { note in

Also, change the Section header above it to:

Section(searchText.isEmpty ? "Notes" : "Matching Notes") {

This will update the label to reflect the filtered list for searches. Run the app, make sure that at least some of your voice notes have a transcript, and enter some search text to see it in action.

Searching the transcripts.
Searching the transcripts.

You can already see the benefit that adding transcripts provides in this app. With those transcripts in place, you can now use Foundation Models to generate much more useful information for the user. You’ll begin doing that in the next section.

Using Apple Foundation Models for Voice Notes

The first question when considering adding Apple Foundation Models or any artificial intelligence features to an app should always be: what value can it provide to the user? Adding AI just to say the app supports it will more likely annoy than impress your users. Always ask where this value is before taking the time to add any feature to an app, and this question is especially important when adding AI.

The first clear feature to add would be a title better fitting the content. Right now, the app bases the default title on the date and time. You’ll now add a feature that lets the app generate a title from the transcript. Create a new Swift file under Services named NoteAnalysisService.swift. Replace the contents of the file with:

import Foundation
import FoundationModels

struct NoteAnalysisService {
  func determineTitle(transcript: String) async throws -> String {
    let trimmedTitle = transcript.trimmingCharacters(in: .whitespacesAndNewlines)

    guard !trimmedTitle.isEmpty else {
      return "" 
    }

    let session = LanguageModelSession()
    let prompt = """
      Analyze the following voice note transcription. Create a concise title of
      a few words.

      Transcription: \(trimmedTitle)
      """
    let response = try await session.respond(to: prompt)
    return response.content
  }
}

Nothing here should look new, as this is the same pattern you’ve used throughout this book when working with Foundation Models. Create a LanguageModelSession, feed it a prompt defining the task, and capture the response.

To use this, return to VoiceNoteStore.swift. Add the following method to update the title of an existing note after the updateTranscript(_:for:) method:

private func updateTitle(_ title: String, for noteID: VoiceNote.ID) {
  guard let index = notes.firstIndex(where: { $0.id == noteID }) else { return }
  notes[index].title = title
  saveNotes()
}

This method will take the id of an existing note and update the title. Now find the stopRecording() method. Then find the Task you added and the await transcribeRecording(note) call. Replace it with:

let noteTranscript = await transcribeRecording(note)
guard let noteTranscript = noteTranscript else { return }
let noteAnalysis = NoteAnalysisService()
let title = try await noteAnalysis.determineTitle(transcript: noteTranscript)
updateTitle(title, for: note.id)

This code now captures the returned transcript in noteTranscript. It then checks to ensure the transcript exists and ends the method with the return if transcribeRecording(_:) returned nil. Next, the code creates an instance of the NoteAnalysisService before calling the just-created determineTitle(transcript:) method with the transcript. It then calls the updateTitle(_:for:) method to set the title for the note.

To see this in action, run the app and record a new note. After a few seconds, you should see the transcript followed by an appropriate title for the note.

Note with model created title.
Note with model created title.

Note the flow here. You intentionally create the note first, then let the transcription run before using Foundation Models to create a title. This produces a more polished result for the user, as they are not held up waiting for steps until the title appears. This is a direct application of Foundation Models that reduces work for the end user: analyzing a transcript to produce a useful title.

Now that you’ve explored using Foundation Models against the transcripts generated through machine learning, feeding one artificial intelligence aspect into another, you can see the power of this feature. While a simple text response would work well for the title and summary, it is less suited to more structured information, such as extracting actionable items from the voice notes. In the next chapter, you’ll continue this app and learn how to use guided generation to produce better results as you expand the data gathered using Foundation Models, and how to present the information to the user.

Conclusion

In this chapter, you’ve taken an existing app and begun integrating Apple Foundation Models into it. The first step was to take the audio data in the app and convert it to text that you can feed into Foundation Models, which you completed in this chapter. This also provided a way to give the user a better experience by allowing them to search these transcripts for specific text. You ended the chapter by feeding the transcript into the model to produce a title for the note based on the contents.

You’ve likely noticed that you spent most of this chapter producing the information for the model, and only at the end used the model and presented the result to the user. That’s not by accident. Proper use of Foundation Models takes more than a prompt and response. You need to produce the data fed to the model and present that response to the user. In the next chapter, you’ll extend the app to produce valuable information using Foundation Models and explore how to present this data to the user.

Key Points

  • A staged pipeline produces a better user experience than waiting for everything to occur. Here, the app creates the note immediately, then transcribes it before generating the title.
  • SpeechTranscriber is the modern framework to perform voice transcription, the process of converting speech to text. This is a vital first step in analyzing the data in the recording using Apple Foundation Models.
  • Using speech transcription requires permissions and inclusion of an appropriate NSSpeechRecognitionUsageDescription key. The first run will generate a permissions prompt just as using the microphone does.
  • The AssetInventory class manages machine learning model downloads. You use it to check the status before use and handle the downloading, supported, unsupported, and @unknown default cases, ensuring the app behaves gracefully across device states.
  • Adding AI without purpose will frustrate and annoy users. Never add AI features for the sake of doing so. Determine how these frameworks can add value to users and make your app more useful.
  • Each element of machine learning provides value, but feeding one framework into another, such as using speech transcription on recorded audio into Apple Foundation Models, can accomplish tasks no one framework can.
Have a technical question? Want to report a bug? You can ask questions and report bugs to the book authors in our official book forum here.
© 2026 Kodeco Inc.