This repository hosts a corpus of climate change related texts by German political parties and MPs, spanning late 2013 - early 2025. It includes the following parties:
- Christlich Demokratische Union (CDU) & Christlich Soziale Union Bayerns (CSU), jointly as CDU/CSU
- Sozialdemokratische Partei Deutschlands (SPD)
- Alternative für Deutschland (AfD)
- BÜNDNIS 90/DIE GRÜNEN
- Die Linke
Texts are paragraphs from parliamentary speeches, social media posts, election manifestos, and press releases by official party sources and MPs. Filtering for the climate change topic is done using the klima-klassifikator (Memminger et al., 2026). Paragraphs are between 10-170 tokens in length. The corpus is annotated with the following metadata:
- author/speaker (if applicable)
- party
- date
- id
- token count
- climate keyword counts
This corpus builds on the AfD-CCC (Stede & Memminger, 2025).
Structure of the corpus
├── /bundestag # Bundestag speech paragraphs
│ ├── afd_bt_climate.csv
│ ├── cducsu_bt_climate.csv
│ ├── greens_bt_climate.csv
│ ├── left_bt_climate.csv
│ └── spd_bt_climate.csv
│
├── /ep # European Parliament speech paragraphs
│
├── /manifestos # Election manifesto paragraphs
│
├── /press # Press release paragraphs
│
├── /telegram # Telegram channel messages
│
└── /twitter # Tweet IDs
DEClimate contains 3,236,030 tokens, over 50,278 climate change related paragraphs. The table below lists the corpus composition in tokens.
| Domain | CDU/CSU | SPD | AfD | Greens | Left | Total |
|---|---|---|---|---|---|---|
| German Parl. | 504,393 | 452,821 | 190,236 | 436,020 | 156,032 | 1,739,502 |
| EU Parl. | 63,086 | 33,554 | 31,470 | 52,858 | 10,334 | 191,302 |
| Twitter(*) | 88,010 | 106,943 | 53,672 | 343,233 | 80,050 | 671,908 |
| Telegram | 20,405 | 1,060 | 10,453 | 2,822 | 5,727 | 40,467 |
| Press | 58,170 | 71,825 | 124,931 | 168,766 | 83,548 | 507,240 |
| Manifestos | 9,253 | 9,326 | 2,726 | 38,091 | 19,660 | 79,056 |
| Total | 743,317 | 675,529 | 413,488 | 1,041,790 | 355,351 | 3,229,475 |
* The Twitter files only contain tweet IDs and not the tweet texts.
- German Parl. speeches: 2013-2021 SpeakGer (Lange and Jentsch, 2023), 2021-2024 GermaParl (Blaette and Leonhardt, 2025), 2024-2025 DIP API
- EU Parl. speeches: 2013-2024 ParlLawSpeech (Schwalbach et al., 2025), 2024-2025 Open Data API
- Twitter: 2016-2021 Lasser et al. (2022), 2024-2025 ???
- Telegram: 2019-2025 scraped with FROG (Primig and Fröschl, 2024)
- Press releases: 2013-2017 PARTYPRESS (Erfort et al., 2023), 2017-2025 webscraping & directly supplied by the Green party
- Manifestos: 2013, 2017, 2019, 2025 via the Manifesto Project (Lehmann et al., 2025) Version 2025-1. See the Manifesto Project Terms of Use in the respective directory.
Texts from PARTYPRESS and the Manifesto Project were included with permission of the dataset creators.
If you use our dataset, please cite the paper:
Ronja Memminger and Manfred Stede. DEClimate: A Dataset of Climate Change Discourse in German Politics. In Proceedings of the Konferenz zur Verarbeitung natürlicher Sprache (KONVENS). Hamburg, Germany, 2026.
If you use the manifestos, also cite the Manifesto Project Version 2025-1 (see below).
DEClimate © 2026 by Ronja Memminger & Manfred Stede is licensed under CC BY-NC-SA 4.0. Commercial usage is prohibited. Appropriate credit must be given.
Andreas Blaette and Christoph Leonhardt. 2025. GermaParl Corpus of Plenary Protocols.
Cornelius Erfort, Lukas F Stoetzer, and Heike Klüver. 2023. The PARTYPRESS Database: A new comparative database of parties’ press releases. Research & Politics, 10(3).
Kai-Robin Lange and Carsten Jentsch. 2023. SpeakGer: A meta-data enriched speech corpus of German state and federal parliaments. In Proceedings of the 3rd Workshop on Computational Linguistics for the Political and Social Sciences, pages 19–28, Ingolstadt, Germany. Association for Computational Lingustics.
Jana Lasser, Segun Taofeek Aroyehun, Almog Simchon, Fabio Carrella, David Garcia, and Stephan Lewandowsky. 2022. Social media sharing of low-quality news sources by political elites. PNAS Nexus, 1(4).
Pola Lehmann, Simon Franzmann, Denise Al-Gaddooa, Tobias Burst, Christoph Ivanusch, Jirka Lewandowski, Sven Regel, Felicia Riethmüller, and Lisa Zehnter. 2025. The Manifesto Data Collection. Manifesto Project (MRG/CMP/MARPOR) (Version 2025-1). Berlin: WZB Berlin Social Science Center/Göttingen: Institute for Democracy Research (If-Dem).
Ronja Memminger, Dietmar Benndorf, and Manfred Stede. 2026. Politische Texte zum Thema Klimawandel automatisch identifizieren: Ein XGBoost Modell. In Book of Abstracts - DHd 2026, pages 608–609.
Florian Primig and Fabian Fröschl. 2024. Introducing the FROG tool for gathering Telegram data. Mobile Media & Communication, 12(2):449–453.
Jan Schwalbach, Lukas Hetzer, Sven-Oliver Proksch, Christian Rauh, and Miklós Sebők. 2025. Parllawspeech. (Version 1.0.0) [Data set]. GESIS, Köln.
Manfred Stede and Ronja Memminger. 2025. AfD- CCC: Analyzing the climate change discourse of a German right-wing political party. In Proceedings of the Fourth Workshop on NLP for Positive Impact (NLP4PI), pages 163–174, Vienna, Austria. Association for Computational Linguistics