Azure Content Safety

Nov 15 2024 · Python 3.12, Microsoft Azure, JupyterLabs

Lesson 02: Text Moderation With Azure Content Safety

Exploring Text Moderation in Content Safety Studio

Episode complete

Play next episode

Next
Transcript

In this segment, you’ll get hands-on experience with how text moderation works in the platform, and you’ll also learn how to add filters and blocklist terms to customize text moderation response.

Create an Azure Account

Before using the Azure Content Safety platform, you’ll need to create an Azure account, if you don’t already have one. You can do that by heading here: Azure Account Creation. Then, click on “Try Azure for Free” as shown on the open page.

You’ll then be redirected to a new page in the new tab, and you’ll be asked to sign in to your Microsoft account. If you have one, you can use it to sign in. If you need to create a new account, then click on the “Create one!” link.

NOTE: This demo will use my existing Microsoft account.

Once you’ve finished signing in, you’ll be shown a form for creating your free Azure account, and you’ll need to share/verify two details to proceed:

  • Your basic information, such as name, contact, address, etc.
  • Your credit card details.

Fill out the form, and you’ll be redirected to the Azure Dashboard. Since I already have an existing Azure account, I will skip this step. You may be asked to complete a small transaction during the credit card verification process. If any charge is deducted, don’t worry—it should be reimbursed within 2-3 working days.

Create a Content Safety Resource

Upon reaching the Azure Dashboard, you’ll next create a Content Safety resource, to get your key and endpoint.

Azure resources allow you to organize and manage your Azure services in a structured way. This enables you to manage permissions, apply policies, and monitor and control costs at the individual resource level.

I’ll be selecting the paid subscription and plan since I’ve already used my free plan. But, you’ll choose Free Trial and Free F0 as subscription and pricing tier respectively.

Start by selecting your subscription and pricing tier. Once selected, you’ll then create a new resource group with a unique name or use an existing one from the available list by clicking on the Resource group dropdown. . Finally, you’ll need to provide a unique resource name in the Instace Details section to successfully create the content safety resource.

Review all other configuration by clicking Next button, and finally click on Create button on the last tab, Review + create.

The resource takes a few minutes to deploy. After it finishes, you can access your newly created resource by selecting the All resources item in the Azure portal and then selecting your resource name from the list.

Congratulations! You’ve successfully set up your Azure account and created a content safety resource to make API requests for moderation needs.

Allocate Permission to Content Safety Studio

Next, you’ll have to provide permission to Content Safety Studio to make an API request to the above created resource. You can do that by clicking Access control (IAM) on the left tab.

Next, click on the Add button. Select Add role assignment from the drop-down that appears. This will open the role assignment page.

Search “cognitive services user” in the text field. Select Cognitive Services User out of all the job function roles that appears. Then click the Next button at the bottom.

You can assign the Cognitive Services User role to different members here. Click on Select members to select and assign the role to yourself. Click the Select button at the bottom of the overlay on the right.

Finally, click the Review + assign button at the bottom to allocate permission to read key content safety resources to make API calls via Content Safety Studio.

Attach Created Content Safety Resource to Studio

You’re now ready to use Content Safety Studio. Click on the Overview button from the left pane. Then, click the Content Safety Studio link inside the Get Started tab. This will open it in a new tab.

Next, click on the settings icon on the top to open the Settings page for the content safety studio. By default, it will open the Resource tab. Under the All resources section, you should be able to see the resource you just created. Select it and then click the Use resource button at the bottom. This will open the studio dashboard once again.

Explore Text Moderation in Content Safety Studio

Let’s start seeing text moderation in action. Click on the Moderate text content card. This opens a page where you can run tests on text and see how the content moderation API performs.

The page is divided into multiple sections:

  • Introduction: This section gives a brief overview of the page’s purpose. It also shares links for sample code and documentation that you can use to get started with Text Moderation API.
  • Try it out: This section contains an acknowledgment that the tests performed here will use your content safety resource.
  • Select a sample or type your own: This section proposes some sample texts that you can use to try Text Moderation API and analyze the results.
  • Test: This is the main section where actual tests are performed. On the left side, it contains a text field where you can add the text you want to analyze. On the right side, it provides options to configure filters and use blocklists to customize the API.
  • View results: Your test result will appear here once the text analysis is finished.
  • Next steps: This section suggests the next steps to follow.

Using Content Safety Studio, you can run two types of tests for text moderation:

  • Run a test on a single text.
  • Run a test on bulk texts.

Testing Single Text for Moderation

From the “Select a sample or type your own” section, select any text on which you would like to run your text moderation API. This will copy the text in the text field below it.

You can then configure the filter to:

  • Enable/disable the category: You can turn on or off the category in which you want the API to analyze/not analyze the text. For example, if you deselect the Sexual category, the text moderation API will only run a check on the text for the remaining categories — violence, hate, and self-harm and will exclude the sexual category.
  • Update the category sensitivity: You can also update the threshold level for each category to allow the level of sensitivity that can be permitted while analyzing the text. Setting the threshold to Low will flag the text as inappropriate for any slightly harmful content. On the other hand, setting the threshold to High will only block seriously dangerous content. For our use case, let’s set the Hate category threshold to Low threshold.
  • Use blocklist: You can further customize moderation by adding terms or phrases that must be blocklisted. The text will be flagged as inappropriate if any of these terms or phrases are referenced in the content. To add a blocklist, click the Add Blocklist button and select Create a New Blocklist from the drop-down menu. In the overlay window that appears, enter a name and description, then click the Create button. Next, click Add Terms to input the terms you want to block and finish by clicking the Done button.

Once you’ve finished your customization, click the Run test button to analyze the text.

The moderation result shows that the selected text is blocked by the moderation API with the reason stated as - Rejected by filter in Violence category

Now, add a phrase from the blocklist terms. Try running the API once again.

The moderation API blocks the updated text again due to the presence of violent material and blocklisted phrases. Finally, try the safe content text example and re-run the test.

This time, the text content is allowed.

Testing Bulk Text for Moderation

Next, you will look at the moderation technique for bulk text. This allows you to test the moderation API result on bulk text and understand moderation API efficacy on your bulk data set, not just single text content.

Similar to a simple test, it also has a few example datasets in CSV format, containing multiple texts and their truth values representing whether the text is safe or harmful. The rest of the sections remain the same as for a single test.

You will use the curated dataset provided in the module for bulk testing. Clone the material repo if you haven’t done so. Then, click on Browse for a file underneath the section Select a sample or upload your own. Navigate to the clone material repo and select the sample-data/post.csv file from the file browse window that appears.

This will load the sample data on food-related social posts relevant to our module’s sample app. Next, click the Run test button to see how the default filter configuration performs on the app’s sample dataset.

Unlike for testing single text moderation, the View result section for Testing bulk text moderation is pretty interesting. Let’s break it down.

First, you have the percentage of allowed content and the percentage that is blocked. In this case, 98.8% of the content is permitted by the content moderation API, while the rest is blocked.

On the right of it, we have metric containing three values:

  • Precision: This measures the accuracy of the moderation API when it classifies a content as harmful. Simply put, it helps answer the question: How much of all the content the model flagged as harmful was actually harmful? Having a precision of 1 in our dataset means that the moderation API was able to flag harmful content as harmful correctly.
  • Recall: This measures the moderation API’s ability to identify all the harmful content. In our case, having a recall value of 0.10 implies that the moderation API could only capture and flag 10% of harmful content as harmful and failed to identify the other 90% of data as harmful.
  • F1 Score: This is the harmonic mean of Precision and Recall metrics. It gives you a balanced view of precision and recall, making it a good overall metric for assessing the effectiveness of moderation APIs in content moderation tasks. In our case, having the F1 score of 0.20 is a sign of a poor score that we can aim to improve.

Finally, at the bottom, you have the Severity detail per record and Severity distribution category graph. These two provide the detailed view of output of each text analysis versus what is expected and severitiy distribution for each category in the dataset.

This shows that the moderation API labeled the majority of the content as safe and a few others as low or medium severity levels for every category.

You can tweak the threshold level for the category to improve the performance metrics. For demo purposes, set the threshold to low for all the categories and re-run the test.

This time, more harmful content was blocked without losing much precision. Yet the Recall value and F1 score remain the same. In the next step, you can add blocklist terms for capturing the text you deem harmful per your platform guidelines.

If the moderation API is not performing well when labeling your dataset, even after configuration tweaks, you should train your own moderation model with your dataset. Training custom models is out of the scope of this module, but if you are interested in it, you can read more about it in Use Custom Categories.

Next, you’ll learn about the content moderation API for text moderation and its implementation in Python so that you can integrate Azure Content Safety into your platform.

See forum comments
Cinema mode Download course materials from Github
Previous: Understanding Text Moderation Using Azure Content Safety Next: Understanding Azure Content Safety Text Moderation API