Skip to main content
Vision-language models on Impala accept video as input. Send it as a reference to an object in S3, or inline as base64. The inference engine splits the video into frames and encodes them with the prompt. You can set the frame rate per request with mm_processor_kwargs:
For S3 access requirements, see Images. The same bucket setup applies.