Instruction 02
Processing images and then drawing boxes around interesting features is a common task. For example, funny-face filters that draw googly eyes need to know where the eyes are. The general workflow is to get the bounding boxes from the observations and then use those to draw an overlay on the image or draw on the image itself.
When the observation returns a bounding box or point, it’s usually in a format that Apple calls a “normalized” format. The values will all be represented as values between 0 and 1. This is so that regardless of your image’s display size, you’ll be able to locate and size the bounding box correctly. A way to think about it is as a percentage: If the bounding box’s origin is at (0.5, 0.5) it’s 50 percent across the face of the image. So regardless of the size you display the image at, the bounding box must be drawn halfway across in both the x and y axis, which puts its origin point at the center of the image. A point of (0, 0) is at the origin point of the image and (1.0, 1.0) will be in the corner opposite the origin. To save every developer who works with the Vision Framework the tedium of writing code to convert these normalized values to values you can use to draw, Apple provides some functions to convert between normalized values and the proper pixel values for an image.
Functions for Converting From Normalized Space to Pixel Space
VNImageRectForNormalizedRect
Converts a normalized bounding box (CGRect with values between 0.0 and 1.0) into a CGRect in the pixel coordinate space of a specific image. Use this when you need to draw a bounding box on the image.
VNImagePointForNormalizedPoint
Converts a normalized CGPoint (with values between 0.0 and 1.0) into a CGPoint in the pixel coordinate space of a specific image. This is useful for translating facial landmark points or other keypoints onto the image.
VNImageSizeForNormalizedSize
Converts a normalized CGSize (with values between 0.0 and 1.0) into a CGSize in the pixel coordinate space of a specific image. This can be used when scaling elements relative to the image size.
Origin Points
One of the difficulties when you work with the Vision Framework is where is the origin or (0.0) point of an image or rectangle? All Vision observations that return coordinates (rectangles, points, etc.) assume that (0.0, 0.0) is the bottom-left of the space. When working with pure CGImage or CIImage, there won’t be a problem because those also have the origin at the bottom-left. However:
Origin Points of Different Image Formats in iOS Development
-
UIImage – Origin Point: Top-Left – Description: High-level image class used for displaying images in iOS.
-
UIView – Origin Point: Top-Left – Description: Fundamental building block for UI elements in iOS.
-
SwiftUI Image – Origin Point: Top-Left – Description: Represents images in SwiftUI.
-
CGImage – Origin Point: Bottom-Left – Description: Low-level image representation in Core Graphics.
-
CALayer – Origin Point: Top-Left – Description: Manages and animates visual content in iOS.
-
CGAffineTransform – Origin Point: Top-Left – Description: Used for performing 2D transformations in UIView contexts.
-
CAAnimation – Origin Point: Top-Left – Description: Animates properties of CALayer objects.
-
SKSpriteNode – Origin Point: Bottom-Left – Description: Used in game development for representing sprites in SpriteKit.
-
MTKView (MetalKit) – Origin Point: Bottom-Left – Description: Displays rendered content from Metal, a low-level graphics API.
-
CIImage – Origin Point: Bottom-Left – Description: Core Image representation of image data.
-
Core Graphics (CGContext) – Origin Point: Bottom-Left – Description: Drawing environment for 2D graphics.
-
SCNNode (SceneKit) – Origin Point: Center – Description: Represents a node in a 3D scene graph.
-
MKOverlayRenderer – Origin Point: Top-Left – Description: Renders overlays on maps in MapKit.
-
PDFPage (PDFKit) – Origin Point: Bottom-Left – Description: Represents a PDF page in PDFKit.
-
CVPixelBuffer – Origin Point: Top-Left – Description: Core Video pixel buffer for managing video pixel data.
-
CMSampleBuffer – Origin Point: Top-Left – Description: Core Media sample buffer encapsulating video and audio data.
-
Front Camera (Portrait) – Origin Point: Top-Left – Description: The coordinate system for the front-facing camera sensor in portrait mode.
-
Back Camera (Portrait) – Origin Point: Top-Left – Description: The coordinate system for the back-facing camera sensor in portrait mode.
Depending on the original format of the image, the pixels might have a different origin point. An image generated with the camera in landscape mode will have a different origin point than one in portrait. Similarly, the underlying data for a UIImage might not be in the “right” orientation for display. For a camera image, there’s EXIF metadata and for UIImage there’s an imageOrientation property so the system knows how to rotate or flip the pixels of the image to look “right-side up” regardless of the actual orientation of the pixels or the camera. This means that if you take a batch of UIImage and push them through your VNImageRequest to get some bounding boxes, the origin point of the bounding boxes and the origin point of the image might not match. So when you go to draw the box on the image, it’ll be in the wrong place, or after you get your image out of your process, it might be rotated.
There are a number of strategies you can use to mitigate this. You might convert your UIImage to a .jpeg because that bakes in the rotation and the image will be the expected orientation. When you’re drawing, you might transform the drawing-space orientation. Another way would be to just apply the same rotation that iOS applies to the pixels before displaying them.
For example, if iOS knows to rotate the pixels 90 degrees clockwise to make the orientation look correct, that means the original pixels are rotated 90 degrees counter-clockwise. So when you’re drawing on the image, you just need to assign the right .imageOrientation value to the final UIImage. You might apply a function like this one:
func convertImageOrientation(_ originalOrientation: UIImage.Orientation)
-> UIImage.Orientation {
switch originalOrientation {
case .up: // 0
return .downMirrored // 5
case .down: // 1
return .upMirrored // 4
case .left: // 2
return .rightMirrored // 7
case .right: // 3
return .leftMirrored // 6
case .upMirrored: // 4
return .down // 1
case .downMirrored: // 5
return .up // 0
case .leftMirrored: // 6
return .right // 3
case .rightMirrored: // 7
return .left // 2
@unknown default:
return originalOrientation
}
}
In the code above, each of the commented values is the rawValue of the .imageOrientation. When you start with a UIImage that has an .up orientation and you convert it to a CGImage and then draw on it in a CGContext, it’ll be .downMirrored when you convert it back to a UIImage. So when creating the final UIImage, just assign it that orientation and iOS will take care of it.
Remember, this is just one way to deal with the origin point rotation problem. Now that you’re aware it’s sometimes an issue, you’ve got a likely culprit to examine when your code isn’t drawing boxes or points where you expect.
Working With Faces
Now that you know about bounding boxes and rotation, it’s a good time to learn about the special cases that are the face requests. Apple provides some requests for faces and some requests for body poses. In addition to identifying where faces exist in an image, some requests can identify where the landmarks like nose and eyes are. Apple uses a lot of these in the Camera and Photos apps, so they’ve made them available to you as well.
iOS 11
- VNDetectFaceRectanglesRequest and VNFaceObservation: Detects faces in an image by finding the bounding boxes of face regions.
- VNDetectFaceLandmarksRequest and VNFaceObservation: Detects facial features such as eyes, nose, and mouth in detected face regions.
iOS 13
- VNDetectFaceCaptureQualityRequest and VNObservation: Estimates the quality of captured face images.
- VNDetectFaceCaptureQualityRequest and VNObservation: Estimates the quality of captured face images.
iOS 14
- VNDetectHumanBodyPoseRequest and VNHumanBodyPoseObservation: Detects and tracks human body poses in images or videos.
- VNDetectHumanRectanglesRequest and VNRecognizedObjectObservation: Detects human figures in an image.
You’ll notice that a few of the requests share a VNFaceObservation return type. However, based on the descriptions, it seems like they might return different kinds of data. They do. Remember that observations are subclasses. One of the parent classes returns the boundingBox of the observation. VNFaceObservations also have optional values for roll, pitch and yaw to help place the orientation of the face and then they have a complex property of landmarks that is of type VNFaceLandmarks2D. This contains a lot of information about where the edges of the eyes are, where the left eye and where the right eye is. There are landmark entries for the pupil in the eye, so you can determine if the eye is open or closed.
These requests follow the same pattern as all the others, so you should have no trouble using them.