Instruction 01
Whether your app is an accessibility tool, data-entry automation or a document scanner, the ability to recognize and work with text opens new possibilities. Apple’s Vision Framework provides a few request types that are specifically designed for these tasks. Here’s a chart for the main text-related Vision request types along with their response observation types and when they were introduced.
iOS 11
- VNDetectTextRectanglesRequest and VNTextObservation: Detects regions of text in an image.
iOS 13
- VNDetectTextRectanglesRequest and VNTextObservation: Performs Optical Character Recognition (OCR) to detect and recognize text.
iOS 15
- VNRecognizeHandwritingRequest and VNRecognizedTextObservation: Recognizes handwritten text in images.
Because these are Vision requests that work with images, building a request and processing the observations should look familiar to you.
let textDetectionRequest = VNDetectTextRectanglesRequest { request, error in
guard let observations = request.results as? [VNTextObservation]
else { return }
// Handle detected text rectangles
}
The handler type is the same one you’ve been using all along as well.
let requestHandler = VNImageRequestHandler(cgImage: inputImage.cgImage!,
options: [:])
do {
try? requestHandler.perform([textDetectionRequest])
} catch {
// Handle the error
}
The only thing that should be new to you is how to handle the two observation types and determining when it’s better to use a particular request type.
Differentiating Between VNDetectTextRectanglesRequest and VNRecognizeTextRequest
When the Vision Framework first arrived, VNDetectTextRectanglesRequest was the only option. This request type gave you the bounding boxes of all the areas in an image that likely contain text. If you wanted to actually perform recognition, you had to use a CoreML model.
Then, a few years later, VNRecognizeTextRequest appeared. The observations from this request type give the bounding boxes of text and also a String of the text itself. So the VNRecognizeTextRequest does all the things that VNDetectTextRectanglesRequest does and more. Why didn’t Apple get rid of the earlier request type?
Possibly, it’s because recognizing just the rectangles of text in an image is a lot less resource intensive and faster than performing the OCR step. Your app might just want to highlight areas of text, to draw the user’s attention to it. When working with video, where speed is the most important consideration, you can show the user the text areas to help them center their camera and then only recognize the text when they’re ready.
Another difference between the two request types is the granularity. By default, VNRecognizeTextRequest provides one observation per line of text. This will give you the bounding box for the entire line of text as well as the String and a confidence score for the entire line of text. The confidence score tells your app how accurate the Vision framework thinks it read the text. If your app wants to preserve paragraph formatting or things like that, you’ll need to have it process the bounding boxes of the lines and look for gaps between them. The VNRecognizedTextObservation has a nifty function, though. You can pass a Range<String.Index> to boundingBox(for:) and it’ll return the bounding box for that range of characters in the line. This lets you highlight each word separately or highlight only certain words in the line, for example.
In demo code you typically see for a VNRecognizeTextRequest, you’ll notice that it works with the topCandidates(1).first, which is Vision’s best guess as to what the text String contains. This is a classification observation, though. As you saw in earlier lessons, there could be many classifications, each with a different confidence score. The topCandidates array is always sorted by decreasing confidence score and the array never contains more than 10 items in the array. By providing a 1 in the brackets, your code is telling the Vision framework that you want only one element returned. This means your app could ask for a few and then compare the topCandidates to see how different the strings are to help you determine accuracy. You could also flag particular lines with low confidence scores to your user, where you’d like them to double-check the recognized string.
When VNDetectTextRectanglesRequest returns, it gives you the bounding box of the entire area of text. That could be one line or multiple lines. You can also get the bounding box for each character in the area of text if you set the reportCharacterBoxes property of the request to true. Then, the returned observation will have an array of bounding boxes in its characterBoxes property. As you can imagine, this will slow things down a bit.
Here’s a comparison of each request type.
The VNDetectTextRectanglesRequest is faster, less resource-intensive because it detects only text areas. It’s ideal for highlighting text areas without OCR and useful in real-time scenarios like video processing. It can return bounding boxes for individual characters with reportCharacterBoxes set to true.
The VNRecognizeTextRequest provides bounding boxes and recognized text (String). It allows precise bounding boxes for specific character ranges within lines. It provides confidence scores for text recognition, aiding accuracy checks. It also offers multiple possible text interpretations through topCandidates.