Vision Framework

Oct 9 2025 · Swift 5, iOS 26, Xcode 26

Lesson 03: Exploring Text Detection & Recognition

Demo

Episode complete

Play next episode

Next
Transcript

Here’s the demo app again. If you want to code along, you can find this version in the Starter folder for the materials for this lesson. When you build and run, you can see there’s a new tab for text recognition. The text recognition will be similar to the face recognition from the previous lesson. After Vision recognizes the text, the user can cycle through the rectangles to see the bounding box in red and also the recognized string for that line.

But right now, as you can see, if you load an image and try to recognize the text, nothing happens.

Open TextDetectionViewModel. By now the pattern should be familiar. One difference with text recognition is this recognitionLevel property. Your choices here are accurate and fast. You can experiment with both options with the kinds of images you’ll process. Fast is less accurate, of course, but it might be what you need if you’re working with live video frames or something like that.

This time, the app is using a VNRecognizeTextRequest here. In the completion handler it’ll expect an array of VNRecognizedTextObservation types. Like you did with the face recognition rectangles, the view model will have an array of objects and the View will let the user cycle through them. In this example, the array is a tuple of the bounding box and the string.

self?.textRectangles = request.results?.compactMap {
  guard let observation = $0 as? VNRecognizedTextObservation,
        let topCandidate = observation.topCandidates(1).first else { return nil }
  return (observation.boundingBox, topCandidate.string)
} ?? []

You can use functional programming techniques to map the array and then grab the observation so you can get the bounding box and the top candidate so you can get the string. Using compactMap instead of map probably isn’t necessary because the observations aren’t supposed to include nil values, but it’s safer.

The handler and the rest of the code are all the same as with the other demos. If you switch to the TextDetectionView, you see the familiar pattern again. However, because of the tuple, the code looks a little different. To get the rectangles, you’re passing in viewModel.currentText?.0, which means the first item in the tuple. Lower, you present the string by using if let (_, text) = viewModel.currentText. The underscore is here to let the compiler and other readers know that you’re only interested in one part of the tuple here.

Build and run the app. Select an image with some text in it and then click the Detect Text button. You see that you’re getting one observation per line, and the text is accurately recognized.

Now, select an image that isn’t as clean. This image has been scanned and printed a bunch of times, with the background getting noisier each time. Look what happens when you Detect Text with this image. The big, hand-written text gets recognized, and most of the other text is recognized, but the word “town” on the last line isn’t showing up. You can try to fix this by pre-processing the image with a filter.

Go back to the TextDetectionViewModel. After creating the request, but before creating the handler, you’ll apply a filter.

The first step is to import the libraries, so at the top add an import for CoreImage and the built-ins.

import CoreImage
import CoreImage.CIFilterBuiltins

Now, you can convert the cgImage variable to a CIImage.

let ciImage = CIImage(cgImage: cgImage)

You don’t need a guard here because the initializer isn’t an optional one. A CGImage type will always convert to a CIImage type.

The first filter to try is an exposure adjustment.

let exposureAdjustFilter = CIFilter.exposureAdjust()
exposureAdjustFilter.inputImage = ciImage
exposureAdjustFilter.ev = 1.0
guard let exposureAdjustedImage = exposureAdjustFilter.outputImage
  else { return }

You put a guard or some other optional check for the outputImage of a filter. If the filter fails for some reason, it won’t output an image. Don’t forget to update the request handler to use its ciImage initializer with the exposureAdjustedImage.

let handler = VNImageRequestHandler(ciImage: exposureAdjustedImage, options: [:])

Also, set a breakpoint here on the line where you create the handler. Now, build and run.

Select the poor-quality image like before and click Detect Text. When the breakpoint hits, you can see what the filter actually did to the image. Find exposureAdjustedImage in the variables window of Xcode. Highlight it and then click this little eye icon at the bottom. This is the quick-look button, and it’ll show you what the filtered image actually looks like.

Remember from before that when you convert your UIImage to a CG or CIImage, it won’t necessarily be in the original orientation. You can also see that it looks a little better, but it’s still kind of dirty. Let the app resume and click through to see if it recognized the word “town”.

It’s still not recognizing the word on the last line correctly. It gets “town” but thinks that the closing quote is an asterisk.

CIFilters can get chained together, so next you can add a contrast filter to try to get the black to be more black and the grays to disappear.

You can look at the documentation to see the allowed values for the .contrast adjustment in this filter and the .ev adjustment in the prior one. You’ll go with a value of 4 for the contrast, which is kind of in the middle and up the exposure to 2.5 since it didn’t look all that brighter at 1.0.

let contrastAdjustFilter = CIFilter.colorControls()
contrastAdjustFilter.inputImage = exposureAdjustedImage
contrastAdjustFilter.contrast = 4
guard let processedImage = contrastAdjustFilter.outputImage else { return }

Finally, don’t forget to pass the processedImage output of the chain of filters to the handler.

let handler = VNImageRequestHandler(ciImage: processedImage, options: [:])

Build and run and load the image. When it hits the breakpoint, use the quick look button to see how the filters affected the processedImage.

That’s a big difference from the original. The gray background noise is almost gone, but the higher exposure didn’t wash out the foreground too much. Let the app resume and see if it recognizes the last line.

Success! There isn’t a perfect setting for any of these filters; for this demo, you just played around with values and this image until you found some that worked with it. You might decide to hard-code values or provide your users with some slider or other control so they can try to clean the image.

If you’re wondering, the filters aren’t so drastic that they ruin good images. Here, you can load the good image from before, and the text gets recognized just as before.

See forum comments
Cinema mode Download course materials from Github
Previous: Instruction 02 Next: Conclusion