Long podcasts, lectures, interviews, and educational videos can contain hours of useful information, but finding the parts that actually matter takes time. Just The Gist is a free Mac application designed to make that process much easier by turning spoken content into clean transcripts and concise summaries directly on your computer.
The app is built around a simple idea: spend less time watching or listening from beginning to end and more time understanding the information you actually need. A two-hour podcast, for example, can be reduced to a short, readable summary with important moments clearly highlighted.
One of its most appealing qualities is the local-first approach. Transcription is handled on the Mac using Whisper, while the built-in summarization model can also run locally. This means users can process their content without automatically sending audio to an online service.
The interface follows the familiar feel of a native Mac application rather than trying to imitate a complicated web dashboard. A library keeps processed videos and audio organized, while users can add new media and move between transcripts, summaries, starred items, and content that is still being processed.
The workflow is particularly straightforward. Add a YouTube link or drop a supported media file into the application, let the local transcription finish, and then review the generated summary. Key topics and timestamps make it easier to jump around the source instead of searching through a long transcript manually.
Transcription uses Whisper locally, with the website demonstrating performance at around 2.8× realtime using Whisper large-v3 on Apple Silicon. Actual speed will naturally depend on the Mac model, recording quality, duration, and language, but the local processing approach can make the experience feel considerably more private and responsive than uploading every recording to a remote service.
The summaries are designed to focus on the substance of a recording rather than simply shortening every sentence. Key topics and important moments are surfaced alongside the transcript, giving users a practical way to understand lengthy material quickly.
The application can handle more than YouTube videos. Users can bring in podcasts, lectures, conference talks, interviews, meeting recordings, and other speech-based media stored on their Mac. Supported formats include MP3, M4A, WAV, MP4, and MOV.
Language support is another useful feature. Transcription and summarization do not have to use the same language. For example, an English lecture can be transcribed in English and summarized in Spanish, while a German podcast could be summarized in English.
For users who prefer larger or different language models, the application also supports optional cloud providers and OpenAI-compatible local servers. This gives users more control over the balance between convenience, model capability, privacy, and hardware usage.
Privacy is one of the strongest aspects of the product. By default, audio transcription takes place directly on the Mac using Whisper, and the built-in summarization model also works locally. The website states that audio does not leave the computer under the default setup.
Cloud processing is optional rather than mandatory. If a user chooses a cloud provider, only the completed transcript is sent for summarization rather than the original audio. This local-first design makes the application particularly interesting for people working with private meetings, research recordings, lectures, interviews, or other material they would rather keep on their own machine.
Pros
Cons
The application is completely free to download and use, with no subscription and no locked features. There is also no requirement to create an account or enter a credit card for the standard local workflow.
The project is supported through voluntary donations. Users who find the application useful can contribute to help fund its continued development and maintenance.
Optional cloud models can be connected when users want to work with services such as OpenAI, Claude, or OpenRouter. The local workflow itself does not require a paid API subscription.
Many transcription and summarization services are built primarily around cloud processing. They can be convenient, but they often require users to upload recordings to external servers and may operate through subscriptions or usage-based pricing.
This application takes a different approach by making local processing the default. That distinction is especially valuable when privacy matters or when someone wants to avoid recurring costs for processing their own recordings.
It also goes beyond simple transcription. The combination of transcripts, concise summaries, key topics, and clickable timestamps creates a workflow that is closer to a personal research assistant for spoken content. At the same time, users who need more powerful cloud models are not locked out, since optional integrations are available.
For Mac users who regularly deal with long videos, podcasts, lectures, meetings, or recorded conversations, this is a remarkably practical approach to turning hours of spoken material into something that can be understood in minutes. The local-first architecture is its biggest differentiator, while support for multiple file formats, multilingual transcription, clickable timestamps, and optional model integrations gives it plenty of flexibility.
The fact that there is no subscription, account requirement, or mandatory API key makes the barrier to trying it extremely low. Someone who simply wants to understand a two-hour podcast without spending two hours listening to it can start with a local workflow and get straight to the important information.
It will not be the right choice for everyone because it is currently built specifically for Apple Silicon Macs, but within that audience it offers a compelling combination of privacy, simplicity, and useful AI-powered summarization.
Yes. The application is free, has no subscription, and does not lock its core features behind a paid plan. Development is supported through optional donations.
No. The default local workflow does not require an account, API key, or credit card. Optional connections to cloud providers are available for users who want them.
No, not in the default setup. Transcription runs locally using Whisper, and the built-in summarization model also runs on the Mac. Cloud processing only occurs if the user deliberately chooses a cloud provider.
The application supports YouTube links as well as MP3, M4A, WAV, MP4, and MOV files.
Yes. Podcasts, lectures, meetings, conference talks, interviews, and other recordings containing speech can be processed and summarized.
Yes. Whisper supports numerous languages, and the language used for the transcript can be different from the language selected for the summary.
Yes. The application can communicate with OpenAI-compatible servers, allowing users to connect local models through platforms such as Ollama and LM Studio.
An Apple Silicon Mac with an M1 chip or newer is required, along with macOS 14 Sonoma or later. Intel Macs are not supported.
Yes. OpenAI and Claude can be connected as optional cloud summarization providers. OpenRouter is also supported.
No. YouTube is supported, but users can also add local audio and video files, making the application useful for podcasts, lectures, meetings, interviews, and personal recordings.
AI Summarizer , AI Transcription , AI Speech to Text , AI Podcast Assistant .
These classifications represent its core capabilities and areas of application. For related tools, explore the linked categories above.
Website unavailable — View Alternatives