Instruction 02
It’d be nice if the only images your app needed to process were dark, black text on a clear, white background. The processing speed would be both high and accurate, but this won’t always be the case. Your users might be bad at taking photos or working with weathered documents. Some fonts are notoriously hard for OCR systems. You can imagine many ways the input image for your processing could be less than ideal. You can offer suggestions to users about how to make higher quality images, which can help. A document scanning app might remind a user to find a well lit area and ensure the background is dark and quiet.
Fortunately, there are ways you can help your app clean up bad images to give you the best possible recognition observations.
Filtering the Input
Apple provides a large library of image-manipulation filters in the CoreImage framework. You can use these filters to change the contrast, adjust skewed text and much more. A deep dive into CoreImage filters is way beyond this lesson’s scope. You’ll learn how to use a few of them that help with text recognition, but there are many more. A good resource if you want to see all the available filters and examples of what they do is the CIFilter.io site. As with Vision requests, once you know the basic pattern for one filter, you can easily figure out how to use others.
Working with the filters introduces a new image type, CIImage, to your workflow. You might remember the entire rotation and mirroring problem from before and how converting your image to CIImage will be difficult. Thankfully, CIImage and CGImage have the same origin point location. So working with both frameworks on the same image is fine. Whew!
The basic workflow is to create a CIImage version of the input image. Apply filters to that image to clean it up. Then, use that image in the request handler by using the ciImage: initializer for the handler. Then, if you’re going to draw rectangles on the image or something like that, use the untouched CGImage version just like you’ve been doing all throughout these lessons.
Some CIFilters to consider for problematic input images are in this list.
- CIColorControls adjusts the brightness, contrast, and saturation of an image, which can improve text visibility by enhancing contrast between text and background.
- CIGaussianBlur applies a Gaussian blur to reduce noise and smooth out an image, which can help in focusing on the text by reducing distractions from background details.
- CIEdgeWork emphasizes the edges of objects within an image, making the boundaries of text more distinct.
- CIEdges detects edges in the image and highlights them, which can enhance the sharpness of text against its background.
- CINoiseReduction reduces noise in the image, making the text stand out more clearly, which is especially useful in low-light or high-noise conditions.
- CISharpenLuminance increases the sharpness of an image by enhancing the luminance channel, making text crisper.
- CIExposureAdjust adjusts the exposure of the image, brightening dark images or correcting overexposed ones.
- CILineOverlay converts the image to a line drawing, emphasizing the text’s structure, which can be beneficial when dealing with low-contrast images.
- CIMinimumComponent reduces an image’s colors to the darkest value present, helping to isolate text in images with complex or colorful backgrounds.
- CIColorInvert inverts the colors of an image, which can make light text on a dark background more detectable by a Vision request.
Filtering with any CIFilter can look like this:
import CIImage
let ciImage = CIImage(cgImage: inputImage.cgImage!)
let filter = CIFilter(name: "CIColorControls")!
filter.setValue(ciImage, forKey: kCIInputImageKey)
filter.setValue(1.0, forKey: kCIInputContrastKey)
let outputImage = filter.outputImage
First, the image is converted to a CIImage. Then, an instance of a filter is created. Each filter has some different parameters to set. All CIFilters have an .inputImage and .outputImage property. There’s no need to execute a separate “process” function for a CIFilter - as soon as the .inputImage property gets set the filter provides an .outputImage. You can apply one or many filters to your image. It isn’t uncommon to see one CIFilter feeding into another CIFilter. After the image has been filtered, use the ciImage: initializer for the handler.
let recognitionRequestHandler = VNImageRequestHandler(ciImage: outputImage,
options: [:])
Because of its history, CIFilter uses “stringly-typed” identifiers. That means, creating a filter involves typing a string of its name. This is prone to error, of course, and the compiler can’t help you see your mistakes. Fortunately, about iOS 14, Apple provided some type-safe initializers and properties for most of the filters called the CIFilterBuiltins.
So to restate the code above using the built-ins, you’d write something like this.
import CoreImage.CIFilterBuiltins
let ciImage = CIImage(cgImage: inputImage.cgImage!)
let filter = CIFilter.colorControls()
filter.contrast = 1.2
filter.saturation = 1.0
let outputImage = filter.outputImage
Now, the compiler can help look for mistakes and the code is easier to read. Here’s a list of filters from before with their type safe names from CIFilterBuiltins when it exists
-
CIColorControls to
CIFilter.colorControls() -
CIGaussianBlur to
CIFilter.gaussianBlur() - CIEdgeWork is not available as a type-safe initializer
-
CIEdges to
CIFilter.edges() -
CINoiseReduction to
CIFilter.noiseReduction() -
CISharpenLuminance to
CIFilter.sharpenLuminance() -
CIExposureAdjust to
CIFilter.exposureAdjust() - CILineOverlay is not available as a type-safe initializer
- CIMinimumComponent is not available as a type-safe initializer
-
CIColorInvert to
CIFilter.colorInvert()
By pre-processing images with CIFilter types, you can make your text recognition more accurate. Unless you’re making a general purpose text recognition app, during development, you should try to get some examples of the kind of images you’ll process so you can figure out which filters will work best for your app.