Voice and Media Settings
Configure speech-to-text, text-to-speech, image and audio understanding, plus image sanitization and the media processing options available.
Overview
Voice and Media settings control how the agent handles voice, audio, image, links, and media content.
Voice Settings
You can configure:
-
1
Text-to-Speech
Allows the agent to speak responses.
-
2
Speech-to-Text
Allows transcription of voice input.
-
3
Talk Mode
Enables continuous voice conversation.
-
4
Voice Wake
Allows activation using a wake command.
Media Settings
You can enable:
-
- Image understanding
- Audio understanding
- Link understanding
- Image generation
Image Sanitization
You can define:
-
- Maximum image dimension
- Maximum image file size
Firecrawl Integration
Firecrawl converts websites into clean, readable content for better parsing.
Click Save after configuration.

