Overview
Llama Guard 4 12B is a content safety classifier that reads text or text-plus-image inputs and returns a clear safe or unsafe verdict, along with the specific policy category that triggered the flag. If you run a platform, build AI-powered tools, or moderate user submissions, getting a reliable second opinion on whether content crosses a line is slow and expensive when done manually. On Picasso IA, Llama Guard 4 12B does that review automatically, returning structured judgments in seconds. It checks for things like hate speech, self-harm content, and graphic violence, so your team can act on clear signals rather than review every piece from scratch.
How It Works
- Write a system prompt that defines the safety policy you want the model to apply, including which violation categories to watch for and how strict the threshold should be.
- Add the text you want evaluated in the prompt field, and optionally include images if you need visual content checked alongside written input.
- Adjust temperature and sampling settings to control response consistency, or leave them at defaults for standard classification behavior.
- Send the request and receive a structured output: a verdict (safe or unsafe) and, when unsafe, the specific category label that applies.
- Route the result into your moderation queue, logging system, or automated workflow to take action immediately.
Frequently Asked Questions
Do I need programming skills or technical knowledge to use this?
No, just open Llama Guard 4 12B on Picasso IA, adjust the settings you want, and hit generate.
What does Llama Guard 4 12B actually output?
It returns a classification verdict: either "safe" or "unsafe." When content is flagged, it also returns the specific violation category, so you know exactly what rule was triggered and can respond accordingly. This makes the output actionable rather than just binary.
Can I check images as well as text?
Yes. The model accepts a list of images alongside your text prompt, letting you evaluate multimodal content in a single request. This is useful for platforms where users post both written content and visual attachments at the same time.
How do I customize which rules the model enforces?
You provide a system prompt that describes the policy the model should apply. You can name specific categories to watch for, set the strictness level, or add any custom guidelines relevant to your community or platform.
How long does a classification take?
Most requests return a verdict within a few seconds. Processing time depends on the length of the input text and the number of images included, but short text-only inputs are typically the fastest.
What happens if I disagree with a classification result?
You can refine the criteria in your system prompt and re-run the request. Rewording the policy description or adjusting the violation thresholds often shifts borderline cases in the direction you expect. Picasso IA lets you iterate as many times as you need without hitting usage caps.
Where can I use the outputs?
The verdict and category label are plain text, so you can paste them into a spreadsheet, feed them into a review queue, or use them as input to another step in an automated content pipeline.