Instruction 01
A Vision Request is a class object that enables you to ask the Vision Framework to perform particular analysis of an image. Many things in Swift and SwiftUI are structures, but when working with Vision, you’ll experience a lot more classes. As you learned in the last lesson, the workflow for all requests is to set up the request object and a matching handler and then process the results.
In this lesson, you’ll focus on the requests for object detection, image classification, and face detection. Recognizing text has some quirks, so it’s covered in the next lesson.
Choosing a Request Type
Apple provides a number of request types. All the request types inherit from the VNRequest class. VNRequest is an abstract class, meaning you never use it directly. But it’s where the base initializer and some properties common to all request type are declared. Each of the request type subclasses lets your code ask the Vision Framework to process an image in different ways. Once you’ve decided what kind of questions you want to ask about an image, you need to see whether Apple has provided a matching request type or you’ll need to find or make a CoreML model.
Here’s a handy list of just some the request and observation types for non-human and non-text requests:
- VNDetectBarcodesRequest and VNBarcodeObservation: Detects barcodes in an image.
- VNDetectRectanglesRequest and VNRectangleObservation: Detects rectangular shapes such as business cards or book covers.
- VNClassifyImageprintRequest and VNFeaturePrintObservation: Generates a print of the image for classification.
- VNCoreMLRequest and VNCoreMLFeatureValueObservation: Uses a Core ML model to perform image analysis tasks.
- VNDetectHorizonRequest and VNHorizonObservation: Detects the horizon in an image.
- VNTrackObjectRequest and VNDetectedObjectObservation: Tracks a specified object across multiple frames.
- VNTrackRectangleRequest and VNRectangleObservation: Tracks a specified rectangle across multiple frames.
- VNDetectSaliencyImageRequest and VNSaliencyImageObservation: Detects contours in an image.
- VNGenerateImageFeaturePrintRequest and VNFeaturePrintObservation: Generates a compact representation of an image’s visual content.
- VNClassifyImageRequest and VNClassificationObservation: Classifies the overall content of an image.
- VNDetectContoursRequest and VNContoursObservation: Detects contours in an image.
- VNGenerateAttentionBasedSaliencyImageRequest and VNSaliencyImageObservation: Generates a saliency map indicating areas of an image likely to draw human attention.
- VNGenerateObjectnessBasedSaliencyImageRequest and VNSaliencyImageObservation: Generates a saliency map highlighting objects in an image.
- VNClassifyJunkRequest and VNClassificationObservation: Identifies and filters out junk or irrelevant content in images.
- VNClassifySceneRequest and VNClassificationObservation: Classifies scenes in an image.
- VNDetectTrajectoriesRequest and VNTrajectoryObservation: Detects and tracks the movement trajectories of objects.
- VNRecognizeAnimalsRequest and VNRecognizedObjectObservation: Detects and recognizes animals in an image.
- VNDetectAnimalRectanglesRequest and VNRecognizedObjectObservation: Detects bounding boxes of animals in an image.
- CalculateImageAestheticsScoresRequest analyzes an image for aesthetically pleasing attributes.
As you can see, there are lots of options, and this is just a subset of what is available. The return observations for each request are matched to that request, so you always need to refer to Apple’s documentation. Some request observations contain a CGRect, a String, or a CGPoint. Some contain some basic data as well as a class that’s specific to the kind of data. It’s important to use the right observation for the request.
Keep in mind as you look at newer documentation, as seen in the CalculateImageAestheticsScoresRequest type, the “VN” prefixes have been removed starting with iOS 18.
Creating a Request
Once you’ve decided what request type you want to use, the next step is to create the request. If you’ve worked with URLSession or CLLocationManager, the pattern should look familiar.
import Vision
let request = VNDetectFaceRectanglesRequest { (request, error) in
guard let observations = request.results as? [VNFaceObservation] else {
print("No faces detected")
return
}
// Process the observations
}
After ensuring the Vision Framework has been imported, create the request and then a completion handler. The request passes in an error if something went wrong and passes in itself. Before it’s executed, the request.results array is nil. If the request is successful, the .results array is populated with an array of the associated observation type. Some request types have configuration settings and some don’t. A request type might have access to different versions of its model or let you suggest cropping. Generally, the code creates the request and then sets any options after it’s created rather than trying to set the options during request initialization. It makes for easier-to-read code.
The error object is of VNErrorCode type, so it has information specific to what might go wrong during request processing. These errors might be that the model couldn’t be loaded or that requested hardware resources failed or that you are requesting an option that doesn’t exist for that request type. As long as you don’t receive the dreaded catch-all .internalError, you should be able to troubleshoot pretty quickly.
Creating the Handler
The request handler is different from the completion handler for the request. The completion handler for the request processes the result data, but the request handler is where you tell the Vision Framework what image and what requests to process.
Whereas the request type passed in an error object to the completion handler, a handler type is a throwing type. Therefore, to catch errors, put it in a do...catch block when you execute it. Among the errors that the handler throws are incorrect image data format, mismatched request types for the handler type, and the ever frustrating .internalError. Below is an example of handler creation.
let handler = VNImageRequestHandler(cgImage: image, options: [:])
do {
try handler.perform([request])
} catch {
print("Failed to perform request: \(error)")
}
The handler takes as its input the image to process as a CGImage and some options. An image request handler actually has a number of different initializers for different use cases. For instance, if you’re working with video frames, they might be CMSampleBuffer or CVPixelBuffer – there are initializers to accept those types. Also, you might have your image as raw Data, and there are ways to initialize with that directly. You might even be working with remote images and can pass in a URL that points at the image. Initializing a request handler with a CGImage or CIImage is the basic way to go, though.
Generally, you can pass in an empty dictionary for the options. Some cases where you might want to pass in options are when you’ve done some preprocessing of the image using a CIContext; you can pass in the context so the handler doesn’t need to make a new one. You can also pass in some “camera intrinsics”, which are things like focal link and the distance between the center of the camera lens and the center of the image. You’d want these when working with 3D models and augmented reality (AR) apps.
You also might notice that the request is passed in as an array. A handler is linked to a single image. So if you want to make multiple requests about that image, you can pass them all in at one time. The completion handlers of each request execute when the handler has processed that request.
Once you’ve created the handler with the image and created some requests, call the .perform([VNRequest]).
Interpreting the Results
Once the handler has processed the request, the completion block of the request executes. The request.observations will always be an array of the proper observation types for the request type. Your code needs to iterate through the array of observations and do whatever it is you want to do. Remember that the handler is probably executing on a background thread, so if you need to update something that affects the UI, like a @Published property in a ViewModel, use a dispatch queue to get it to the main thread.
The type of observation determines what data you’ve given to interpret. Detection observations tend to provide a bounding box of the detected thing or the center point of an item. Classification observations return a string label of the classified object and a confidence score.
Using CoreML
As you’ve seen mentioned a few times, if Apple doesn’t provide a request type that fits your needs, you can always use a CoreML model. Working with a CoreML model requires only a few changes to your code. Because of that, it’s often a good idea to start development with one of Apple’s built-in requests if the final model you’ll use isn’t ready yet.
To use a CoreML model, drag it into your Xcode project just as you would with some media files or image files. Then, you instantiate the model and use it to instantiate a request. Don’t forget to import CoreML. The code below creates a model using Resnet50, which is a commonly used image-classification model.
import CoreML
guard let model = try? VNCoreMLModel(for: Resnet50().model) else {
fatalError("Failed to load ResNet50 model.")
}
The next step is to create a request using the model and then process the results in a completion handler as normal.
let request = VNCoreMLRequest(model: model) { request, error in
DispatchQueue.main.async {
if let results = request.results as? [VNClassificationObservation] {
// Sort and filter results by confidence
}
}
}
The handler code is the same as with built-in requests. Also, remember that the request handler code takes requests as an array, so you could certainly give it a mixture of Apple-provided requests and your CoreML requests.
The documentation for a CoreML model tells you what kind of observations it returns. Because different models answer different questions about the image, the observation types are different. Here’s a list showing the different types:
- VNClassificationObservation: Used with models that return classification labels and confidence scores such as ResNet50, MobileNet and InceptionV3.
- VNCoreMLFeatureValueObservation: Used with models that output a feature value, often a multi-dimensional array (e.g., feature vectors) such as Autoencoders, feature extractors like MobileNetV2 (used in transfer learning) and custom models that output embeddings.
- VNRecognizedObjectObservation: Used with object detection models that return bounding boxes, class labels, and confidence scores for each detected object such as YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector) and Faster R-CNN.
- VNPixelBufferObservation: Used with models that output an image or image-like data as a pixel buffer (e.g., image generation or segmentation masks) such as GANs (Generative Adversarial Networks) for image generation and segmentation models like DeepLab.
- VNImageAlignmentObservation: Used with models that perform image registration tasks or processes that align images, often in scenarios like panorama stitching or image registration.
- VNContoursObservation: Used with custom contour detection models and edge detection models or processes that detect contours or outlines in images.
- VNFeaturePrintObservation: Used with image similarity models or models used in biometric recognition (e.g., facial recognition) that generate a feature print or fingerprint for comparison tasks, often used in image similarity detection.
- VNSaliencyImageObservation: Used with saliency detection models amd custom attention-based models that identify the most salient or attention-grabbing parts of an image.
If it isn’t clear from the model’s documentation what observation type it returns, you can usually see the type by using Xcode’s model viewer or run a request and inspect the return type to see what class it actually returns.