Instruction 01

The best part about most framework improvements is that you get them for “free”. Your current code just works better. Observations are more accurate and come back faster. Some of the observations provide details they didn’t provide before.

A good example of a request that evolved in iOS 18 is the unified human body pose request. Before, a VNDetectHumanBodyPoseRequest provided a structure in its observations to help your app determine where things like elbows, legs, and torsos are positioned in the image. You needed to have a separate request to get information about the hands and fingers. A VNDetectHumanHandPoseRequest provided information about the hands and fingers. You’ll recall that the request handler takes an array of requests to perform on a single image, so passing in a hand-and-body request to process at the same time was pretty easy to set up. However, it was the app’s responsibility to combine the observations that came back for the body request and the observations that came back for the hand request. Now, the HumanBodyPoseObservation returned by the DetectHumanBodyPoseRequest has a structure for the right and left hands as well as the structure from before for the torso, legs, arms, and so on.

Aesthetics and Finding the Best Image

Though many request types improved for iOS 18, only one was completely new: CalculateImageAestheticsScoresRequest, which returns an array of ImageAestheticsScoresObservation objects. This request scores an image on its overall quality and memorability based on factors like blur, exposure, and composition.

With the new aesthetics scores and updated attention-based saliency requests, Apple is helping apps determine what is the “best” image out of a set and what is the most important thing in an image.

In iOS 13, Apple gave developers a VNDetectFaceCaptureQualityRequest that worked for faces. That request type determined “best” by evaluating lighting, whether the face was in the middle, and other image elements. This new request type expands that further to look at images to determine if they’re “good”. Additionally, the VNImageAestheticsScoresObservation returns an isUtility Boolean for images that are in focus with good lighting but aren’t of an interesting subject. Screenshots would be a good example of an image that would get a true value for isUtility.

Sample images with their aesthetics score and whether or not Vision considers them to be utility images. That this awesome portrait of a dog only garnered a 0.7 rating is outrageous though!
Sample images with their aesthetics score and whether or not Vision considers them to be utility images. That this awesome portrait of a dog only garnered a 0.7 rating is outrageous though!

As with many range values that Apple gives apps, the aesthetic score for an image is a Float between 0.0 and 1.0. So when the Vision framework calculates aesthetic scores, the ones with the highest score are the “best”. Remembering back to earlier lessons: the first step is to create a request that tells your app what question you’re asking about the image and what to do with the observations. The request handler is what brings the image into the process and executing the .perform method on the handler is how you bring them together.

Until now, every example you’ve seen has one request, one image, and one request handler. When working with batches of images, things change slightly. A request handler takes a single image. So you’ll likely create a single request because you want to perform the same request on each image. But you’ll create as many request handlers as you have images to process. So far, this seems reasonable. However, when the request’s closure returns with observations, there’s no clear way to associate the observations back to the original image. A developer needs to match the score to the image manually.

Some code might help further explain. Consider this snippet that uses a structure to help keep the observation associated with the image:

struct ProcessedResult {
  let image: CGImage
  let observations: [VNObservation]
}

let images = <some array of cgImages>
var processedImages: [ProcessedResult] = []

for image in images {
  let handler = VNImageRequestHandler(cgImage: image, options: [:])
  let request = VNCalculateImageAestheticsScoreRequest{ result, error in
  guard let observations = request.results as?
    [VNImageAestheticsScoresObservation], error == nil else {
      //either got no observations or got an error
      return
  }

    processedImages.append(ProcessedResult(image: image,
      observations: observations))
  }

  do {
    try handler.perform([request])
  } catch {
    //do something with the error
  }
}

What’s this code doing? By creating an array of ProcessedResult, it’s possible to keep the observations associated with their image. This method isn’t the only way to solve the problem, but it avoids trying to keep two arrays in sync or making dictionaries. There isn’t a retain-cycle problem with using image in the request completion handler because the cgImage is a value type. The same goes for processedImages because arrays in Swift are value types. This method is possibly prone to race conditions, though, because two of the completion handlers might try to append their data at the same time. So you could add some serial processing queues or embrace the new async/await pattern you’ll learn about in the next section.

See forum comments
Download course materials from Github
Previous: Introduction Next: Instruction 02