KIEZSign in
4 min readArtur Arslanov

Every source behind this English guide to Berlin is in German

All 116 source URLs behind 221 place records are tagged as German. Not one English-language source. That reads at first like a gap in the research, and it is the opposite of one.

There is a number in this corpus that stopped us when it came out of the database, because the first reading of it is wrong.

Every place record in the research corpus carries at least one source URL. There are 116 distinct ones across 221 records. Every single one is tagged de.

A hero figure showing all 116 source URLs are in German, above a bar chart of the publishers behind them

Source language across 116 distinct URLs. Full size.

A hundred per cent of anything usually means a bug. The obvious suspicion is that the language field defaults to de and nothing ever sets it otherwise, so the first thing to do is look at what the URLs actually are. They are berlin.de's district pages, Tagesspiegel, Berliner Zeitung, B.Z., Mit Vergnügen, Weddingweiser, the venues' own sites, and three Reddit threads. There is no Time Out Berlin, no Culture Trip, no listicle from a travel publisher in London, no aggregator.

Why this is the good version

An English-language guide to Berlin has two ways to get its facts. It can read what other English-language guides to Berlin have written, which is the cheap option and produces a document that agrees with every other document and inherits all their mistakes at once. Or it can read what Berlin publishes about itself, which is slower, requires German, and produces a document that occasionally disagrees with the consensus because it went to the source.

This corpus took the second route, apparently without anyone deciding to. The tier breakdown makes the shape clear: 49 official sources, 36 from the city press, 11 from local press, 17 listicles and three with no tier at all. Two thirds of the evidence is either a public body or a Berlin newspaper.

The three Reddit threads are worth naming rather than hiding, since they are the weakest thing in the set. They are tagged with the lowest tier, they support soft claims about atmosphere rather than facts about hours or prices, and they are the sort of source a reader is entitled to discount.

Who the publishers actually are

Every source carries a tier, assigned when it was recorded. There are five, and they rank roughly by how much a claim from that source is worth without a second one behind it.

Tier Sources What it means
official 49 A public body, or the venue's own site
city_press 36 Tagesspiegel, Berliner Zeitung, B.Z.
listicle 17 Mit Vergnügen and similar round-ups
local_press 11 Weddingweiser, district blogs
none 3 The Reddit threads

By publisher, Berlin.de is the largest single source with 26 URLs, then Tagesspiegel with 19, Berliner Zeitung with 12 and Mit Vergnügen with nine. Everything after that is a long tail of ones and twos: a venue's own site, a district blog, a Berliner Rundfunk piece.

Weighted by records rather than by URLs the balance shifts further towards official sources, because the berlin.de borough index pages each back many places at once. On that count it is 135 official, 36 city press, 35 listicle, 11 local press and four from Reddit, out of 221.

Where the page-level citations point

The 116 URLs are the source of the place records. The district pages carry their own citations, and those are a much larger and more concentrated set: 972 inline references across 101 distinct domains.

A ranked bar chart of the domains behind 972 citations, with berlin.de far ahead of everything else

Domains behind 972 citations on the district pages. Full size.

berlin.de alone accounts for 238 of them, which is roughly a quarter of every reference on the site. visitberlin.de adds 52. After that it falls away fast: planetarium.berlin has 32, schloss-gutshof-britz.de has 26, Tagesspiegel 20, mauerpark.berlin 20.

A long tail of 101 domains sounds diverse until you notice the top two are the city's own portal and the city's own tourist board, and together they carry 30% of everything. That is a defensible concentration for factual claims about opening times and public sites, and a bad one for judgement. The city's tourist board is not a neutral party on the question of whether somewhere is worth visiting.

The cost of doing it this way

Three of the sixteen district research files came back written in German, because the run followed its sources into the language they were in. Those files are translated alongside as separate English versions rather than being overwritten, so the original stays intact and the count of what the research produced stays honest.

The larger cost is verification. Reading a German municipal page correctly is harder than reading an English summary of it, and a mistranslation is invisible on the page in exactly the way a wrong opening time is. We run a check that every external URL in a published post traces back to a URL in the research corpus, which catches a link written from memory. It cannot catch a German sentence read wrongly, and nothing automatic can.

If you want the other half of this picture, the note on the five pages behind a third of the guide covers how concentrated those 116 URLs are, which is the more uncomfortable number of the two.

MethodSourcesData