# Mediagenda open data

Daily aggregates per outlet, on Europe/Amsterdam dates. One CSV per table per month.
No article text is included; every number can be traced to the articles behind it on the site.

## agg_daily_outlet-YYYY-MM.csv
day, outlet, articles (first seen that day), news_articles (label-eligible news), words, snapshots (front page
snapshots taken), front_page_items (placements recorded), labeled (articles with a judgement).

## agg_daily_outlet_key-YYYY-MM.csv
day, outlet, dimension, key, articles (distinct articles), weight (meaning depends on the dimension).
Dimensions: theme (website category; weight = rule score), topic and primary_topic (judge), event_type, stance, tone,
gap (headline versus body), content_kind, frame, issue (judge main issue), actor_portrayal (key = actor|portrayal),
quoted_source, entity and entity_title (computed mentions; weight = mention count), lexicon_theme and lexicon_term
(word choice; weight = hits), hint (rule-based event hints), wire (wire service marks), section, kind, hour and weekday
(publication time), front_minutes_theme (minutes on the front page, 15 per snapshot), front_list (front page list names).

## agg_daily_outlet_metric-YYYY-MM.csv
day, outlet, metric, value (average over that day's news articles), n (articles averaged).
Metrics: title_sentiment and body_sentiment (local model, -1 to 1, tone not bias), title_words, question_share,
emotive_share, clickbait_share, superlative_share, flesch_douma (readability), quotes_per_article, speaker_share,
wire_share, words, headline_edit_share, publish_lag_minutes.

Method versions are listed in meta.json on the site. Licence: CC BY 4.0 for these aggregates.
