Used this tool? Rate it
Used this tool? Rate it
ScreenApp is an AI screen recorder and analysis platform that treats a screen capture, a meeting, or a voice memo as the raw input to a text pipeline rather than as a finished artifact. You record or upload, and the service returns a transcript, a chaptered summary, speaker labels, extracted action items, and a searchable index you can question later. The home page states the ambition in four words — Your Window Into Your Recordings — and follows it with a blunter promise: stop watching, start understanding. The product is operated by ScreenApp Pty Ltd, a company registered in Sydney, New South Wales, Australia, and its interface currently ships in fourteen languages.
The pivot behind that positioning is documented outside the company's own marketing. In a customer story published by the inference provider Groq, ScreenApp is described as a product that began as a simple screen recorder before its founders noticed users did not merely want to record — they wanted to use what was inside the recordings. The Australian technology outlet Startup Daily, publishing an Antler investor memo, dates the company's formation to 2022, names the founders as Andre Smith and Buddhika Jayawardhana, and records that at pre-seed stage it had 17,500 active monthly users and had processed more than 130,000 videos in a single month.
What separates ScreenApp from the crowded AI-notetaker field is less any single feature than the breadth of the surface area: eighteen distinct tool entry points sit on top of one capture-and-transcribe core, ranging from a meeting recorder and note taker to a clip maker, a compressor, a subtitle generator and a text-to-speech engine. Whether that breadth is a strength or a dilution depends on what you need, and this page tries to give you enough verified detail to decide.
Recording is not confined to one entry point. ScreenApp ships a web app, a Chrome extension, native desktop applications for macOS and Windows, and mobile applications for iOS and Android. The desktop builds matter for workflows that need system audio capture, larger uploads or background recording — capabilities a browser tab cannot reliably provide. Alongside live capture you can upload existing files or import from a URL, though both of those paths are gated behind a paid plan.
The most technically candid part of the official site is its accuracy page, which explains that this speech to text layer does not bet on a single vendor. Jobs are routed by source platform, length, channel layout and language across three primary providers: Whisper Large-v3 running on Groq infrastructure for breadth (the widest language coverage and the fastest path for long-form audio), Google Gemini 3.1 Flash Lite for short audio under five minutes, and xAI Grok Speech-to-Text for phone calls and multi-channel recordings. Cloudflare Workers AI, Fireworks AI, Mistral and Baseten sit behind those as fallbacks so that a single vendor outage does not fail a job.
The same page is unusually direct about the layer above transcription: summarization, chat and AI analysis run on Google Gemini end to end, and the page states plainly that the product is not powered by GPT-4, ChatGPT or Claude. For anyone evaluating tools on the basis of which model reads their meetings, that is a rare and useful disclosure.
ScreenApp advertises 99 languages for transcription, but only 25 of them get word-level speaker diarization through xAI Grok STT. Every other language is transcribed as text without per-speaker attribution. If your work depends on knowing who said what — depositions, panel discussions, multi-party interviews — the supported list narrows considerably, and it is worth checking your working language against it before committing.
Rather than a single marketing accuracy figure, the accuracy page publishes a matrix across three recording conditions. English measures 4.2% word error rate in studio conditions, 7.8% in a conference room, and 12.4% in the field on a handheld phone microphone. Japanese, Korean and Mandarin land at 19.8%, 19.2% and 20.4% respectively in field conditions. The methodology is stated: eighteen hours of public-domain audio per language, scored with the jiwer library, retested quarterly, with punctuation and capitalization excluded from penalties. Numbers this specific are checkable, which is the point of publishing them.
Beyond raw text, the pipeline segments a recording into chapters with their own summaries, identifies speakers, extracts action items and decisions, and formats output into reusable templates. The Groq case study describes exactly this shape of output — meeting minutes, interview transcripts, polished reports — and it is what distinguishes the product from a transcription service that hands you an undifferentiated wall of text.
The analyzer handles material that is not purely spoken. It examines video frame by frame to detect scenes, objects and key moments, identifies sound quality problems and speakers in audio, runs object detection and text recognition on images, and extracts tables and key information from PDFs and documents. This is where the video analysis quota on the pricing page applies, and it is metered separately from ordinary transcription.
Once material is indexed you can question it. Chat with Recordings works within a single file; Ask AI Across Files searches your whole library and is withheld from the free tier. Search itself is listed as always free and unlimited, as are voice dictation, document refinement and export.
One implementation detail deserves attention because it changes what the accuracy numbers mean. When a YouTube video already publishes a caption track, ScreenApp reads that caption track instead of transcribing the audio — no speech recognition runs at all. That path is fast (median 14.7 seconds against 41.1 seconds previously) but the text is YouTube's, not ScreenApp's, so none of the published word error rates describe it. The company states this limitation itself, and also notes the before-window sample was only 61 imports, placing the true median speedup somewhere between 1.9x and 3.7x.
The note taker page builds its entire pitch on an absence: No bot joins your calls. Recording happens on your own device, so participants see no third-party attendee announcing itself. ScreenApp names Otter, Fireflies and Fathom directly as tools that do announce themselves, and targets consultants, therapists, lawyers and sales teams who cannot have a visible bot in sensitive conversations. This is a genuine architectural difference rather than a cosmetic one, though it shifts the disclosure obligation squarely onto you — more on that in the limitations section.
For internal meetings the value is simpler arithmetic: someone currently takes notes, and that person is not fully participating. Automatic chaptering with per-chapter summaries and an extracted action list replaces the after-the-fact write-up, which is the practical shape AI meeting notes take here. Growth and higher plans also include a meeting bot for calendar-joined calls, which sits alongside rather than replaces the device-side recording path.
Students recording lectures gain a searchable index of the semester rather than a folder of unplayed audio. The chaptering matters more here than the transcript: a two-hour lecture with chapter markers is navigable, whereas a two-hour transcript is not. The iOS listing is categorized under Productivity with a 4+ age rating, and the mobile apps are the practical capture surface for an in-person room.
Journalists and qualitative researchers doing meeting transcription need accurate attribution more than they need speed. Here the 25-language diarization list matters, as does the field-condition word error rate — a handheld phone in a noisy room is exactly the 12.4% English case, not the 4.2% studio case. Budget review time accordingly.
Recorded calls become a corpus you can query. The analyzer's speaker identification and the cross-file Ask AI feature let a manager ask questions across many calls rather than replaying each one, and the xAI Grok routing is specifically described as strongest on phone-call and multi-channel audio.
For anyone processing YouTube material — competitive research, course review, content repurposing — the URL import path plus summarization turns the product into a video summarizer that converts hours of footage into readable notes. Just remember which path ran: caption import for captioned videos, full transcription for the rest.
Screen recordings of a workflow, transcribed and chaptered, become draft documentation. The clip maker and subtitle generator finish the job for material that will be shared as video rather than read as text.
Start on the web app if you are recording a browser tab or a quick screen capture, install the Chrome extension if you record from the browser habitually, or install the macOS or Windows build if you need system audio, background recording or large uploads. For in-person conversations, the iOS or Android app is the only sensible choice.
Live recording covers screen, webcam and in-person audio on every plan. Uploading a file and importing from a URL both require a paid plan — on the free tier those two options are marked unavailable, which is a meaningful constraint if your intended use is summarizing existing videos rather than making new recordings.
Transcription is routed automatically; you do not choose a provider. A sixty-minute meeting is described as completing in roughly three minutes end to end, including summarization and chaptering. That figure comes from the vendor and is not independently audited, though the underlying speed change is corroborated by the Groq case study, where a transcription that once took 20 minutes now finishes in about 15 seconds.
The intended workflow inverts the usual one. Open the chaptered summary and the action list first, then jump into the transcript only where a chapter suggests something needs checking. Reading a full transcript defeats the purpose of the tool.
Use Chat with Recordings for a single file. If you are on a paid plan, Ask AI Across Files lets you query the whole library — useful when you know something was said but not in which meeting.
Download of video and audio, PDF export and general export are all paid-tier capabilities. Plan for this: material recorded on the free tier lives inside the product and cannot be taken out.
The single most common misreading of this product's pricing is the time unit. Growth's 600 AI credits and 600 transcriptions are per year, not per month, and its video analysis allowance is 36 per year. Budget accordingly before assuming a plan is generous.
Match the microphone to the accuracy you need. The gap between 4.2% and 12.4% word error rate in English is entirely a recording-conditions gap. A lavalier or dedicated microphone in a treated room produces materially cleaner text than a handheld phone in a café, and no model choice compensates for that.
Check your language against the diarization list before you rely on speaker labels. Transcription reaches 99 languages; word-level speaker attribution reaches 25. Discovering the difference after recording a four-person panel is an expensive way to learn it.
Know which path ran on a YouTube import. If the video had captions, you received YouTube's text, whose quality is entirely outside ScreenApp's control and unrelated to its benchmarks. For anything you will quote, verify against the audio.
Treat the free tier as an evaluation, not a workflow. Three recordings in total, three AI generations per month, and a single transcription per month, with no upload, no URL import and no export, is enough to judge output quality and nothing more.
Set your own calendar reminder if you start the trial. The seven-day Growth trial converts to an annual charge. This is the single most complained-about aspect of the product on independent review platforms, and the mechanism is entirely avoidable with a reminder set on day one.
Record locally when the content is sensitive. The help centre documents a local storage option with cloud sync disabled. For material that should never leave a device, that setting exists and is worth finding before the first sensitive recording rather than after.
Tell people they are being recorded. Because no bot announces itself, the disclosure obligation is yours alone. In many jurisdictions consent is a legal requirement rather than a courtesy, and the invisibility that makes the product attractive is precisely what makes this your responsibility.
Use chapters as the navigation layer. For any recording over thirty minutes, the chapter list is the artifact you will actually use week to week. Skim it first; open the transcript only where a chapter title raises a question.
Client-facing professionals under discretion constraints — consultants, lawyers, therapists and sales teams — are the audience the note taker page addresses explicitly, and the bot-free architecture is a real answer to a real objection.
Students and lecturers get an affordable searchable archive, though the free tier's single monthly transcription means a genuine semester-long habit requires a paid plan.
Small teams needing meeting minutes without a dedicated scribe are well served by the Growth tier, provided their volume fits inside 600 transcriptions per year — roughly two per working day.
Researchers and journalists benefit from the published per-language error rates, which let them estimate review effort honestly instead of trusting a single marketing accuracy figure.
Enterprise buyers with SSO, API and compliance requirements are served by the Enterprise tier, which starts at $199 per month and is the only tier carrying SAML SSO, custom integrations and dedicated support.
Who should look elsewhere: anyone who needs guaranteed refundability, since the terms state fees are non-refundable; anyone whose primary need is heavy video analysis, which stays metered at 120 per year even on the Business tier; and anyone working in a language outside the 25-language diarization list who depends on speaker attribution.
ScreenApp runs on the web and ships six additional clients. All figures below were verified first-hand on 24 August 2026 and stores change daily.
iOS — listed as ScreenApp: AI Voice Recorder, published by ScreenApp Pty Ltd, rated 4.0 stars across 111 ratings. Current version 1.4.47, released 17 August 2026; first published 18 March 2025. Requires iOS 15.0 or later, rated 4+, free to download.
Android — the Play listing carries 3.88 stars across 761 ratings with more than 100,000 downloads, categorized under Productivity with an Everyone content rating. This is the largest independent sample of the four channels checked.
Chrome extension — 4.1 stars across 71 ratings and roughly 20,000 users, carrying a Featured badge, with Google noting the publisher has no history of violations. Note that a widget embedded on one of ScreenApp's own feature pages advertises 4.7 stars from 2.1k ratings for this same extension, which the live store listing does not support; the store figure is the one to trust.
macOS — distributed as a direct DMG download from the official site rather than through the Mac App Store. The build is a universal binary for Apple Silicon and Intel, published on a rolling release with a constant filename, so there is no version string to compare against.
Windows — a native desktop application, downloadable from the site.
Interface languages — the site publishes localized pages in fourteen languages including Spanish, French, German, Italian, Portuguese, Japanese, Korean, Russian, Indonesian, Chinese, Turkish, Dutch and Vietnamese. This is separate from, and much narrower than, the 99 languages supported for transcription.
Four tiers are listed. The figures below are from the official pricing page as read on 24 August 2026; pricing pages change, so confirm against the live page before purchasing.
Free — $0. Three recordings in total (not per month), three AI generations per month, one transcription per month, full transcript included, no video analysis, no file upload, no URL import, no download or export. Enough to evaluate output quality.
Growth — $19 per month billed annually. Unlimited recordings, 600 AI credits per year and 600 transcriptions per year, 36 video analyses per year, meeting bot included, download and export enabled. Sold with a seven-day free trial, after which the page states plainly: 7 days free, then $228 per year.
Business — $34 per month billed annually. Unlimited recordings, unlimited AI credits and transcriptions, API access, white label, custom vocabulary and bulk upload. Video analysis remains capped at 120 per year despite the unlimited framing elsewhere on the tier.
Enterprise — from $199 per month, priced on team size. Unlimited everything, SAML SSO, custom integrations and dedicated support. SSO is available on no other tier.
Two things to read carefully. First, the page headline reads "No credit card required, cancel anytime," while the same page's FAQ states that the seven-day trial does require a credit card and converts automatically. Both statements are on the same page; the first describes the free tier, the second describes the trial. Second, the terms and conditions state that all fees for the Service are non-refundable, which is the policy that governs if a charge lands unexpectedly.
Otter.ai, Fireflies.ai and Fathom are the three competitors ScreenApp names itself, and it differentiates on one axis: those tools join a call as a visible participant, while ScreenApp records from your device. If a visible bot is acceptable or even desirable as an implicit recording notice, those tools are mature alternatives with deep calendar integration.
Loom is addressed directly in the official FAQ, which frames the comparison as $18 per month with a five-minute free limit and AI as a paid add-on, against ScreenApp's $19 annual-billed tier with AI included. Loom remains the stronger choice for polished asynchronous video messaging; ScreenApp is aimed at extracting text from recordings rather than at video presentation.
Dedicated transcription services such as those built directly on Whisper or on Deepgram, AssemblyAI and ElevenLabs — all named in ScreenApp's own benchmark comparisons — will suit anyone who wants raw transcription with API control and no surrounding product.
General-purpose assistants are explicitly ruled out by the analyzer page, which argues that text-based chat interfaces cannot process, watch or listen to video and audio files at all. That is a fair statement of why a media-processing tool exists as a separate category.
Native platform recorders — macOS screen capture, Windows Game Bar, built-in phone voice memos — remain free and adequate whenever you need the file itself and not the text derived from it.
The independent review picture is poor, and the pattern is specific. On Trustpilot the company scores 1.3 out of 5 across 52 reviews, with 94 percent of them one star. The profile was claimed by the company in August 2022, but Trustpilot notes it has no recent history of inviting customers and has not replied to negative reviews. The top mentions cluster tightly: cancellation, subscription, refund, payment, customer communications. Named reviews describe a seven-day trial converting to a $228 annual charge without a clear reminder, and refund requests being declined by reference to the cancellation policy. That account is consistent with the vendor's own published terms, which state fees are non-refundable, and with the pricing page's own trial wording — which is why it reads as policy rather than as isolated mishap. Set a reminder, or do not start the trial.
Store ratings and Trustpilot disagree sharply, and both samples are real. iOS sits at 4.0 across 111 ratings and Android at 3.88 across 761 — mid-range but not poor. These are not reconcilable into one number, and they should not be: app store samples come from general users rating a product experience, while Trustpilot's sample here has self-selected around billing disputes. Read them as answering different questions. Capterra's 5.0 comes from just four reviews and is too small a sample to support any quality conclusion at all.
Quota units are easy to misread. Growth's 600 AI credits and 600 transcriptions are annual figures presented next to a monthly price. Business is described as unlimited but still caps video analysis at 120 per year.
Several official statements contradict each other. The help centre article on safety says the company was founded in 2020; the site's own structured data says 2023; the Antler investor memo published by Startup Daily says 2022. Three different founding years across the company's own materials and an investor document. The same safety article claims all data is stored on ScreenApp's own servers rather than third-party platforms, which sits badly against the security page's acknowledgement that infrastructure runs on AWS and the privacy policy's explicit list of five third-party AI providers. The terms of service still describe plans named Free, Standard and Premium with limits that no longer match the current four-tier pricing page. None of these are fatal, but collectively they indicate documentation that has not kept pace with the product.
Self-reported user counts do not agree with each other. The home page says 8.1 million people and, elsewhere on the same page, 8,159,301+, while the reviews page says 3M+ active users. Both are vendor-reported and unaudited. The reviews page also reports a 4.9/5 average, an NPS of 72 and 94% satisfaction, none of which is independently verifiable and all of which sits in obvious tension with the Trustpilot distribution.
Some published accuracy numbers are projections. The iPhone microphone column in the word error rate table is, in the company's own words, a projection, not a measurement, computed by multiplying field conditions by 1.2. Crediting the company for saying so is fair; treating the number as measured is not.
A cited source has disappeared. The accuracy page attributes its speed figures to a Groq customer story, but that URL now redirects to Groq's blog index and the case study is no longer published at its original address. The content was recoverable only through a web archive. The figures themselves — 20x faster, 15x cheaper, 30% conversion lift, churn halved, ARR up 405% — are vendor statements quoted in a supplier's marketing material, never independently audited, and should be read as such.
No major independent IT press coverage exists. Searches of the significant technology outlets returned no independent review or reported article on this product. The third-party evidence available is a supplier case study, an investor memo published as partner content by an Australian startup outlet, and app store data. That is thinner than a mature category leader would have, and it is stated here plainly rather than padded with affiliate blog posts.
Recording carries obligations the product deliberately does not handle. Because no bot announces itself, nothing in the workflow notifies other participants. Consent requirements vary by jurisdiction and are frequently mandatory. The product gives you invisibility; the legal exposure that comes with it is yours.
No. The privacy policy states directly that ScreenApp does not use your recordings or content to train AI models, and that its third-party AI providers process data only to deliver the requested service and are not authorized to use it for training. It further states that the models are pretrained before deployment and uploaded data is processed solely for service provision. This is an unusually clear statement for a recording product and is the single most important line in that document.
Five are named in the privacy policy, with the data each receives specified: Google, Groq, Mistral, Fireworks AI and Cloudflare. Google (Gemini and Vertex AI) receives transcript text and audio for summarization, chaptering, titling, speaker diarization and translation. Groq receives transcript text and audio for AI chat and transcription. Mistral receives transcript text for AI chat. Fireworks AI and Cloudflare receive audio as transcription fallbacks. Naming sub-processors this explicitly is rare and worth crediting.
All personal data is stored in the United States, which is worth noting given the company is Australian. Data is retained only as long as needed for the stated purposes or as required by law. If you request account deletion through settings, all personal data is permanently deleted within 30 days, with exceptions where retention is legally required. You can ask support for details of anything retained.
Three recordings in total, three AI generations per month and one transcription per month, with the full transcript included. Video analysis, file upload, URL import, download and export are all unavailable. It is sufficient to test output quality on your own material and insufficient as an ongoing workflow.
The trial applies to the Growth annual plan and, per the pricing page FAQ, requires a credit card and converts automatically at the end of seven days to $228 per year. On refunds the terms are unambiguous: all fees for the service are non-refundable. This combination is the source of most negative independent reviews. If you start the trial, set a reminder for day six.
The security page states ScreenApp has achieved SOC 2 Type II certification, independently verified by third-party auditors, and links to a trust centre hosted at trust.inc. That trust centre page confirms SOC 2 Type 2 compliance and lists one framework and 35 controls but publishes neither the audit window nor the auditor. So the claim is corroborated by a third-party trust platform, while the underlying report details are not public. Enterprise buyers should request the report directly.
Not on the device-recording path, which is the product's headline differentiator: recording happens on your own machine and other participants see no additional attendee. A meeting bot is separately available and included from the Growth tier upward for calendar-joined calls. Because the default path is silent, informing participants that you are recording is entirely your responsibility and is a legal requirement in many jurisdictions.
It depends almost entirely on recording conditions. The published English figures are 4.2% word error rate in studio conditions, 7.8% in a conference room and 12.4% in the field on a handheld phone microphone. Japanese, Korean and Mandarin sit near 20% in field conditions. Testing used eighteen hours of public-domain audio per language scored with jiwer, retested quarterly. Note the iPhone microphone column is an estimate rather than a measurement.
Transcription covers 99 languages through Whisper Large-v3. Speaker diarization — attaching a speaker label to each word — covers only 25 of those through xAI Grok STT. Other languages produce text without per-speaker attribution. Separately, the website interface itself is localized into fourteen languages, which is a different and much smaller list.
Partly. Business offers unlimited recordings, AI credits and transcriptions, but video analysis is still capped at 120 per year. Growth is not unlimited at all on the AI side: its 600 credits and 600 transcriptions are annual allowances shown beside a monthly price, alongside 36 video analyses per year. Read the time unit rather than the adjective.