There’s a starter project in the materials repository for this lesson. Here’s the project, which uses the photos picker from the last lesson. If you want to swap out the camera or the Document picker, you should be able to just bring those ViewModels into this project. You’ll need to update the toolbar code here. And change the value for the photoPickerViewModel. Working with the photos picker on the simulator is pretty easy, so that’s what the demo uses.
Build and run this application. You can see it has two tabs and there’s a button at the top right to choose an image from the simulator’s photo roll. If you’ve never worked with the photo roll on the simulator, you can just drag and drop images into it like this. If you have an iCloud account, you can use it to log into the simulator if you want to grab your images and documents that way.
As you can see, nothing happens on the tabs when I select Detect Faces or Classify Image. This demo will make the Classify Image button work, and the next demo will make the Detect Faces button work.
Quit the app and open the project navigator. As with the last lesson, a lot of the boilerplate code, like for displaying a tab bar or making a button, isn’t really the focus of this lesson. So those things are already done for you.
Here in the ObjectDetectionView, there’s a regular SwiftUI View and an ObjectDetectionViewModel. The Vision Framework code goes in the ViewModel. The view is just a VStack that displays an image if it has one and below that displays some text and the button. The Vision requests in this demo all return text about what’s in the image.
Over here in the ViewModel, you can see that the skeleton is already set up. The request is a VNRecognizeAnimalsRequest. Below, the code tells the simulator to process the request on the CPU. If you don’t tell the simulator to use the CPU, Vision requests fail. Even if your Mac has a GPU, the simulator won’t use it, so you’ll need this piece of code.
With every request you need a handler. You can see here the VNImageRequestHandler that takes the image formatted as a CGImage. Scroll back up to the request completion block and add some code to process the results. First, you’ll want to get the results out of the request object and cast them to the appropriate type.
Over here, the documentation is open and you can see that for a VNRecognizeAnimalsRequest you expect an array of VNRecognizedObjectObservations. Clicking that, you can see that each observation will contain an array of labels. These labels are all the objects the model knows about and it will assign a value to each one as to its confidence that the object is in the image.
Now that you know what type of objects you’ll get, you can add the code to get them ready to process.
if let results = request.results as?
[VNRecognizedObjectObservation] {
}
Every label that the model knows about will get a confidence score. The labels that have the highest confidence score are what the model thinks the image contains. You can sort and map the results array to get just the result the model is most confident about.
let sortedResults = results
.sorted(by: { $0.confidence > $1.confidence })
.map { "\($0.labels.first?.identifier ?? "Unknown" )
- \((Int($0.confidence * 100)))%" }
Now, you can populate the string in the ViewModel with the new value and since it’s published to the view, it’ll update the view.
if let topResult = sortedResults.first {
self?.classification = topResult
} else {
self?.classification = "Unknown"
}
Build and run. Navigate to the animals tab and click the button to select an image. Here’s an image of a dog. To see if the model can figure out it’s a dog, click the Classify Image button. The model is 63 percent confident that this is a dog! Selecting another image, see how it does with this one: Unknown.
That’s pretty clearly a horse. For any model that does classification, you can query the model to see what it knows about. Return to the code and uncomment this code. Request objects that do classification have a .supportedIdentifiers function to tell you all the things they can find. Build and run the app again and try to classify the horse. It’s still unknown, but look down in the console: The only animals this request can classify are cats and dogs. Apple’s VNRecognizeAnimalsRequest is built for learning purposes, not for actually classifying animals.
Fortunately, Apple does provide some other requests.
Replace the VNRecognizeAnimalsRequest with a more general VNClassifyImageRequest. Now, replace the VNRecognizedObjectObservation with a VNClassificationObservation. Checking the documentation, you can see that instead of an array of label that is just an identifier property.The observations will have the .confidence like before. Because .confidence is a property of the inherited VNObservation class. So you can sort the results and then create a block of text because the classification observation usually lists many things that are in an image.
Filter the results so you have only things where the model has a confidence score greater than 0.01.
.filter { $0.confidence > 0.01 }
Now, the sort by confidence is the same, but generate the text differently by mapping each result and then joining them into a giant String.
.map { "\($0.identifier) - \((Int($0.confidence * 100)))%" }
.joined(separator: ", ")
Because sortedResults isn’t an array anymore, update the classification property.
if !sortedResults.isEmpty {
self?.classification = sortedResults
} else {
self?.classification = "Unknown"
}
Now, update the supportedIdentifiers code or comment it out again. If you want to update it, you can use code like this:
if let objects = try? request.supportedIdentifiers() {
for object in objects {
logger.debug("Object: \(object)")
}
}
Now, you’re ready to build and run again. Navigate to the animals tab and choose a new image. Now, click “Classify image”. This still doesn’t look right. Looking at the console, you can see that it knows things like a horse and a person. Try a different image. Something is wrong.
It turns out that this model when run on a simulator will always return this as the observations. You can confirm this by running on a device. Here’s a physical device attached to this computer. When you run the app, you can find that same picture of a boy and his dog. Now when you tap “Classify image”, you get sensible results.