How our pipeline works
Every episode we monitor moves through the same four stages: ingestion, transcription, analysis, reporting. Fully automated, continuously running.
Fetching the audio, at scale
Our crawlers continuously poll podcast RSS feeds, hosting CDNs and ad-serving endpoints for the shows under monitoring. New and updated episodes are fetched automatically, usually within minutes of publication.
- Scheduled fetching from RSS feeds, CDNs and ad servers
- Change detection: re-fetch when an episode's audio is replaced
- Format normalization to a common internal representation
- Original files retained for audit and re-analysis
Every word, with a timestamp
Speech-to-text turns each episode into a timecoded transcript. Timestamps are what make verification possible: they let us say not just that an ad was read, but exactly when and for how long.
- High-accuracy speech-to-text with word-level timestamps
- Automatic language detection across markets
- Speaker diarization: who is speaking, host or guest
- Ad-segment boundaries identified in the transcript
Four analyses, one pass
With the audio and transcript in hand, the analysis stage runs every check a campaign needs in a single pass over the episode.
- Ad detection and transcript-to-script matching for verification
- Audio quality metrics: loudness (LUFS), clipping, silence, encoding
- Content classification of transcripts for brand safety
- Availability and accessibility checks across platforms
Evidence, delivered where you work
Results become useful when they reach the right people. Verification outcomes, quality metrics and suitability scores are available the moment analysis completes.
- Campaign dashboards with episode-level drill-down
- Scheduled delivery and discrepancy reports
- API and webhook delivery for programmatic integration
- Exportable evidence trails for billing reconciliation
Built for scale and auditability
Queue-based processing
Every stage is decoupled behind durable queues, so publication spikes never drop episodes — they just take a few extra minutes.
Redundant monitoring
Availability checks run from multiple regions and providers, so a single vantage point never decides whether an episode counts as reachable.
Retained evidence
Original audio, transcripts and analysis results are versioned and retained under a defined data-retention policy, so every reported result can be re-derived and audited.
Want to see the pipeline on your own campaign?
We can run a sample verification on a live campaign and walk you through the results.