Understanding Multi-Modal Content Moderation
Multi-modal content moderation is the approach to moderate content in different formats — such as text, image, video, audio, or a combination of these formats — within a single framework.
Traditionally, content moderation techniques relied primarily on simple algorithms and human moderators. While these methods served their purpose at the time, they now have more obvious limitations. As digital platforms evolved, new content forms like images emerged.
The increased use of these diverse content types, often combined with one another, creates a demand for more complex and advanced moderation systems — this is where the new approach to multi-modal content moderation comes to light.
Here, instead of dealing with text-based and visual-based content, moderation is done separately. Multi-modal systems are meant to analyze them together and decide whether the content is safe. This approach should also improve the accuracy of the overall moderation system.
For example, a social media post may contain offensive text, paired with inappropriate images. A multi-modal moderation system solution will better evaluate both the image and text elements of posts together.
Understanding Multi-Modal Content Moderation for the Fooder app
As you’ll already know by now, the sample app used in this lesson, Fooder, is a social media app for recipes — users can post food recipes and, alongside this, see recipe photos, views, and comments over posts.
Your task for this lesson will be to integrate a robust moderation system that can both analyze images and test them, and make our famous Fooder app a secure and healthy place to check-in. Sounds like a role with heavy responsibility, but you got this!
Here’s what the moderation system will look like. When the user uploads a post (with or without images or photos) and comments, both the text content and image will go through the moderation system. The text will be sent to the text analysis API, and the image will be sent to the image analysis API. The response from both APIs will then be collected and analyzed to ensure the content requested for publishing is safe and as per the community guidelines.
If the content is found to be okay, it’s allowed to be published; if it’s rejected, the user shares details for the reason for rejection and suggests addressing the concerns for publishing. That’s pretty much it.
Let’s start with implementation…