OpenAI Transcribe Skill

Official OpenAI skill for audio-file transcription, speaker separation, and reviewable text output.

The OpenAI Transcribe Skill is an official workflow instruction from the curated openai/skills repository. It is intended for compatible coding agents and harnesses that need to prepare dependable text from existing audio or video files. According to the provider, the workflow first collects file paths, the requested response format, possible language hints, and references for known speakers. It then selects an appropriate transcription route, saves the result, and checks text quality, speaker labels, and segment boundaries. The primary source is the current skills/.curated/transcribe directory in the official OpenAI repository. OpenAI’s official file-transcription documentation adds the API and model context needed to understand the workflow.

A distinct purpose

Transcribe is more than a generic reminder that speech can be converted to text. It defines a repeatable routine for a clear task: turning a completed recording into text, optionally with speaker segments. This separates it from the already catalogued OpenAI Speech Skill, which turns text into spoken audio. Transcribe works in the opposite direction. That boundary keeps the catalog entry useful even when individual model names or client invocation details change. It is not an audio player, an archive system, or a finished transcription interface. Instead, it structures the work of an agent that can use the official OpenAI interface and the bundled local workflow.

Workflow and selection

Inputs are clarified before processing begins. They include one or more recordings, the expected response format, an optional language expectation, and terms whose spelling matters. For ordinary fast transcription, the source recommends a suitable transcription route with a plain text response. When several people need to be distinguished, the specialized speaker-diarization route is appropriate. According to the provider, that variant is designed to identify changing speakers and is not the default choice for every file. Longer audio requires an appropriate automatic chunking or voice-activity configuration. Known-speaker references can help map labels, but they do not replace human review.

Quality, privacy, and boundaries

A saved text file is not automatically a verified record of a conversation. Responsible operators should check names, specialist terms, numbers, accents, translations, and unclear passages against the recording. Speaker labels and time segments deserve special attention when people overlap, background noise is present, or several languages are used. The official documentation describes supported audio formats and a file-size limit; those details can change and should be checked again before production use. The skill does not promise legal evidentiary value, complete accessibility, or error-free recognition.

Audio may contain personal, confidential, or copyrighted material. Before uploading, teams must decide whether sending the material to the selected model provider is permitted, who may access originals and transcripts, and how long both should be retained. Credentials remain in approved local configuration and never belong in skill text, source code, prompts, or logs. Temporary outputs should be limited to necessary content, protected from unauthorized access, and removed according to internal rules. An agent or Codex may create additional logs, so the actual data path of the chosen environment must be documented separately.

Practical use and responsibility

Teams can use Transcribe for interviews, meetings, support recordings, research notes, subtitle drafts, and internal analysis. Repeatable inputs, separate output locations, and a final listening or reading review improve traceability. For sensitive or public releases, subject-matter approval, consent, rights clearance, and editorial responsibility remain with the operator. The license information and repository state should be checked again in the official source before redistribution. This makes the skill a distinct, durable workflow without misrepresenting it as a complete product or a guarantee of transcription quality.

Free
Provider
OpenAI
License
Apache-2.0
Last reviewed
09.09.2026

Repository and documentation

Categories

Compatible with

Claude Code Codex Cursor