Technology

How our pipeline works

Every episode we monitor moves through the same four stages: ingestion, transcription, analysis, reporting. Fully automated, continuously running.

Stage 01 · Ingestion

Fetching the audio, at scale

Our crawlers continuously poll podcast RSS feeds, hosting CDNs and ad-serving endpoints for the shows under monitoring. New and updated episodes are fetched automatically, usually within minutes of publication.

  • Scheduled fetching from RSS feeds, CDNs and ad servers
  • Change detection: re-fetch when an episode's audio is replaced
  • Format normalization to a common internal representation
  • Original files retained for audit and re-analysis
Stage 02 · Transcription

Every word, with a timestamp

Speech-to-text turns each episode into a timecoded transcript. Timestamps are what make verification possible: they let us say not just that an ad was read, but exactly when and for how long.

  • High-accuracy speech-to-text with word-level timestamps
  • Automatic language detection across markets
  • Speaker diarization: who is speaking, host or guest
  • Ad-segment boundaries identified in the transcript
Stage 03 · Analysis

Four analyses, one pass

With the audio and transcript in hand, the analysis stage runs every check a campaign needs in a single pass over the episode.

  • Ad detection and transcript-to-script matching for verification
  • Audio quality metrics: loudness (LUFS), clipping, silence, encoding
  • Content classification of transcripts for brand safety
  • Availability and accessibility checks across platforms
Stage 04 · Reporting

Evidence, delivered where you work

Results become useful when they reach the right people. Verification outcomes, quality metrics and suitability scores are available the moment analysis completes.

  • Campaign dashboards with episode-level drill-down
  • Scheduled delivery and discrepancy reports
  • API and webhook delivery for programmatic integration
  • Exportable evidence trails for billing reconciliation
Reliability

Built for scale and auditability

Queue-based processing

Every stage is decoupled behind durable queues, so publication spikes never drop episodes — they just take a few extra minutes.

Redundant monitoring

Availability checks run from multiple regions and providers, so a single vantage point never decides whether an episode counts as reachable.

Retained evidence

Original audio, transcripts and analysis results are versioned and retained under a defined data-retention policy, so every reported result can be re-derived and audited.

Want to see the pipeline on your own campaign?

We can run a sample verification on a live campaign and walk you through the results.

Request a walkthrough