advancis Artifical Intelligence
SENTINEL
AI Sentinel supports risk detection for content compliance, sensitive content, and prompt injection attacks in two scenarios: model input content and generated content. Additionally, the console provides comprehensive functions such as online testing, data reports, and result query.
Check item confirguration This feature supports setting of appropriate check items for different scenarios and configuring switches for fine-grained labels.
Content compliance detection Detects baseline risks such as politically sensitive content, pornography, and violence, along with violations such as abuse, bias, and harmful values in LLM input or generated content.
Sensitive content detection Automatically identifies, classifies, and grades personal sensitive information and enterprise sensitive information in LLM generated content.
Prompt injection attack detection Identifies deliberately generated violating content in LLM output that bypasses security policies through prompt manipulation (such as inductive or adversarial prompt construction) or technical means (such as encoding obfuscation or multi-round conversation disguise).

Word library management and matchingWhen performing content compliance detection, you can set up lists of risky prohibited keywords or keywords that need to be filtered out before text detection, and then configure detection rules for keyword matching if you need to customize private moderation rules.

Answer library management You can use this feature when performing content compliance detection if you need to replace blocked content with pre-set answers from an answer library.

Online testingThis feature supports online testing of content compliance detection, sensitive content detection, and prompt injection attack detection covered by AI Sentinel to quickly verify the effectiveness of moderation policies.

Result queryThis feature allows users to view the moderation results and returned parameters of checked content through the result query function, meeting requirements such as analyzing high-frequency risk types.

Risk reportsIt allows users to understand the Content Moderation, Sensitive Data Detection, and Prompt Injection Attack Detection call Trend Statistics and call Risk Distribution by viewing risk reports.

Custom detection agentsAI Sentinel lets you configure and manage custom detection agents. This feature uses large language models (LLMs) to let you define custom detection rules, which helps you quickly detect and filter content based on your business-specific categories. This topic describes how to use the custom detection agent feature.