Cosmic Transcriber
Cosmic Transcriber
Soh’s reviewed Windows 10/11 application for long MP3 lectures, interviews, podcasts, Dharma talks, folders, and multi-part recordings. Features resumable batch transcription, TXT/Markdown output, and two complementary OpenAI model modes.
Key Capabilities at a Glance
Engineered to handle long-form audio archives effortlessly, Cosmic Transcriber streamlines batch transcription with robust session state management and output flexibility.
📂 Batch Audio Processing
Drag and drop individual MP3 files or entire nested folder trees for hands-off queueing.
⚡ Resumable Checkpoints
Completed chunks are saved dynamically. Resume interrupted sessions without re-processing.
📝 Dual Export Formats
Outputs pristine plain text (.txt) or clean, formatted Markdown (.md) files automatically.
🔐 Secure Key Management
Optionally encrypt your API key locally using native Windows DPAPI user protection.
Examples of Use & Ideal Workflows
Cosmic Transcriber is tailored specifically for long-form, complex speech content where accuracy and context preservation matter most:
Dharma Talks & Spiritual Lectures
Perfect for multi-hour retreats, oral teachings, and technical Buddhist philosophy. Pair with gpt-transcribe to preserve specialized Pali/Sanskrit terminology and delicate negations while significantly minimizing hallucination risks.
Podcasts & Panel Interviews
Process multi-speaker recordings using gpt-4o-transcribe-diarize to automatically segment dialogue into distinct speaker turns accompanied by timestamps.
Multi-Part Audio Series
Load entire multi-part series or nested folder hierarchies in a single step. The app handles long-file chunking and tracks completion checkpoints file by file.
Two-Pass Editorial Pipeline
Combine the raw verbal fidelity of gpt-transcribe with the structural turn labels of gpt-4o-transcribe-diarize in ChatGPT to generate publication-ready transcripts.
Choose the Model by Purpose
Cosmic Transcriber supports two complementary OpenAI model modes. Selecting the appropriate model for your recording type ensures the best balance between wording fidelity and speaker structure.
gpt-transcribe (Default)
The newer high-accuracy general model. Use it as the controlling text when exact wording, specialist terminology, multilingual speech, and negations matter.
- Supports language hints, literal keywords, and recording context.
- Provides maximum wording fidelity and term accuracy.
- Note: Does not add speaker labels in the app’s current file workflow.
gpt-4o-transcribe-diarize
The older specialised speaker-labelled model. It identifies speaker turns and can include timestamps throughout the audio.
- Identifies individual speaker turns (Speaker A, B, etc.) and timing.
- Note: Cosmic Transcriber does not send prompt/context or structured keyword hints in this mode.
- Note: Speaker letters may reset between independently processed Parts.
Soh’s Recommended Two-Pass Method
Transcribe the same recording once with gpt-transcribe for the most reliable wording and once with gpt-4o-transcribe-diarize for speaker structure. Upload both outputs to ChatGPT; tell it to keep the first transcript as the wording authority and use the second only for speaker labels, timestamps, and turn boundaries. Ask it to mark uncertain alignments instead of guessing.
Copy-Paste Prompt for ChatGPT:
Get an OpenAI API Key
- Sign in to the OpenAI API key page and create a new secret key.
- Set up API billing in the OpenAI Platform billing page. API usage is billed separately from a ChatGPT subscription.
- Copy the key once and keep it private. Do not post it, place it in public files, or share it with other people.
Quick Use Guide
- Download and extract the ZIP, then start Cosmic Transcriber normally rather than as administrator.
- Enter the API key for the session, load it from
OPENAI_API_KEY, or optionally save it for the current Windows user using Windows DPAPI. - Add MP3 files or folders, choose the model and output options, then start transcription. Compatible completed chunks can resume from checkpoints.
0 Responses