There weren’t a large number of improvements to Vision in iOS 26, but there is a new API that extends on the API seen in the last demo. The RecognizeDocumentsRequest API improves on the functionality seen in RecognizeTextRequest by providing the ability to read structural elements of a page, like tables, lists, and important information like phone numbers. It also provides support in 26 languages. The new API can also provide an estimate of camera lens smudge detection, which is actually used in the Camera app in iOS 26. Let’s take a look at the new structure detection in a demo.
Here is the project in the Starter folder in the materials for this lesson. This project is already set to have a minimum iOS target of 26.0, which is different from how the last demo started. In fact, this starter project is the final project from the last demo.
For this demo, you’ll update the app’s text-recognition feature to use RecognizeDocumentsRequest.
Open TextDetectionViewModel.
Change the textDetectionRequest to be a RecognizeDocumentsRequest:
let textDetectionRequest = RecognizeDocumentsRequest()
Remove the next 2 blocks of code that set the recognition level and the special code to set the compute device to the CPU. This API does not need those additions.
Further down the page, where the handler is initialized, update it for this type of request. This type of request is different from the ones encountered to this point. You can invoke the perform method directly on the request and get the observations from the request in return. No handler necessary!
Replace the code in the do-catch block with the following:
// Perform the request on the image data and return the results.
let observations = try await textDetectionRequest.perform(on: UIImage(ciImage: processedImage).pngData()!)
// Get the first observation
guard let document = observations.first?.document else {
throw RecognitionError.documentNotAvailable
}
for paragraph in document.paragraphs {
textRectangles.append((paragraph.boundingRegion.boundingBox.cgRect, paragraph.transcript))
}
var detectedEmail = ""
var detectedPhone = ""
var detectedAddress = ""
//loop through the observations
for result in observations {
for detectedData in result.document.text.detectedData
{
switch detectedData.match.details {
case .emailAddress(let email):
detectedEmail = email.emailAddress
print("detected email \(detectedEmail)")
case .phoneNumber(let phoneNumber):
detectedPhone = phoneNumber.phoneNumber
print("detected phone number \(detectedPhone)")
case .postalAddress(let address):
detectedAddress = address.fullAddress
print("detected address \(detectedAddress)")
default:
break
}
}
}
The code here walks through some of the elements of the returned DocumentObservation type, in this case, the paragraphs and text properties of the document. The paragraphs are another text container, showing that the observation structure is very hierarchical. The code is simply looping through the available paragraphs to show what is in the document.
The text property is a container as well, and here, the code is checking the detectedData property to look for special elements like email addresses, phone numbers, and addresses.
To handle errors, be sure to add the error type near the top of the file:
enum RecognitionError: Error {
case documentNotAvailable
case noElementExists
}
Before trying this out on the simulator, change the exposure value to 1.0, and the contrast to 1.5 - those will work better with the sample image.
Build and run and the code on the simulator, and load the sample wedding invite image from the starter folder. Choose detect text, and as before, you can walk through the paragraphs identified in the image. Take a look at the console though - it has detected an address - even if it is a partial one - down near the bottom of the image.
This example shows that the document reading capabilities of Vision are getting more and more powerful each year. Who knows what it will be able to do in a future pair of Apple Glasses??!