Google Flu Trends (GFT) was a web service operated by Google. It provided estimates of influenza activity for more than 25 countries. By aggregating Google Search queries, it attempted to make accurate predictions about flu activity. This project was first launched in 2008 by Google.org to help predict outbreaks of flu.^[1]

Google Flu Trends stopped publishing current estimates on 9 August 2015. Historical estimates are still available for download, and current data are offered for declared research purposes.^[2]

History

The idea behind Google Flu Trends was that, by monitoring millions of users’ health tracking behaviors online, the large number of Google search queries gathered can be analyzed to reveal if there is the presence of flu-like illness in a population. Google Flu Trends compared these findings to a historic baseline level of influenza activity for its corresponding region and then reports the activity level as either minimal, low, moderate, high, or intense. These estimates have been generally consistent with conventional surveillance data collected by health agencies, both nationally and regionally.

Roni Zeiger helped develop Google Flu Trends.^[3]

Methods

Google Flu Trends was described as using the following method to gather information about flu trends.^[4]^[5]

First, a time series is computed for about 50 million common queries entered weekly within the United States from 2003 to 2008. A query's time series is computed separately for each state and normalized into a fraction by dividing the number of each query by the number of all queries in that state. By identifying the IP address associated with each search, the state in which this query was entered can be determined.

A linear model is used to compute the log-odds of Influenza-like illness (ILI) physician visit and the log-odds of ILI-related search query:

\operatorname {logit} (P)=\beta _{0}+\beta _{1}\times \operatorname {logit} (Q)+\epsilon

P is the percentage of ILI physician visit and Q is the ILI-related query fraction computed in previous steps. β₀ is the intercept and β₁ is the coefficient, while ε is the error term.^{[citation needed]}

Each of the 50 million queries is tested as Q to see if the result computed from a single query could match the actual history ILI data obtained from the U.S. Centers for Disease Control and Prevention (CDC). This process produces a list of top queries which gives the most accurate predictions of CDC ILI data when using the linear model. Then the top 45 queries are chosen because, when aggregated together, these queries fit the history data the most accurately. Using the sum of top 45 ILI-related queries, the linear model is fitted to the weekly ILI data between 2003 and 2007 so that the coefficient can be gained. Finally, the trained model is used to predict flu outbreak across all regions in the United States.

This algorithm has been subsequently revised by Google, partially in response to concerns about accuracy, and attempts to replicate its results have suggested that the algorithm developers "felt an unarticulated need to cloak the actual search terms identified".^[6]

Privacy concerns

Google Flu Trends tries to avoid privacy violations by only aggregating millions of anonymous search queries, without identifying individuals that performed the search.^[1]^[7] Their search log contains the IP address of the user, which could be used to trace back to the region where the search query is originally submitted. Google runs programs on computers to access and calculate the data, so no human is involved in the process. Google also implemented the policy to anonymize IP address in their search logs after 9 months.^[8]

However, Google Flu Trends has raised privacy concerns among some privacy groups. Electronic Privacy Information Center and Patient Privacy Rights sent a letter to Eric Schmidt in 2008, then the CEO of Google.^[9] They conceded that the use of user-generated data could support public health effort in significant ways, but expressed their worries that "user-specific investigations could be compelled, even over Google's objection, by court order or Presidential authority".

Impact

An initial motivation for GFT was that being able to identify disease activity early and respond quickly could reduce the impact of seasonal and pandemic influenza. One report was that Google Flu Trends was able to predict regional outbreaks of flu up to 10 days before they were reported by the CDC (Centers for Disease Control and Prevention).^[10]

In the 2009 flu pandemic Google Flu Trends tracked information about flu in the United States.^[11] In February 2010, the CDC identified influenza cases spiking in the mid-Atlantic region of the United States. However, Google's data of search queries about flu symptoms was able to show that same spike two weeks prior to the CDC report being released.^{[citation needed]}

“The earlier the warning, the earlier prevention and control measures can be put in place, and this could prevent cases of influenza,” said Dr. Lyn Finelli, lead for surveillance at the influenza division of the CDC. “From 5 to 20 percent of the nation's population contract the flu each year, leading to roughly 36,000 deaths on average.” ^[10]

Google Flu Trends is an example of collective intelligence that can be used to identify trends and calculate predictions. The data amassed by search engines is significantly insightful because the search queries represent people's unfiltered wants and needs. “This seems like a really clever way of using data that is created unintentionally by the users of Google to see patterns in the world that would otherwise be invisible,” said Thomas W. Malone, a professor at the Sloan School of Management at MIT. “I think we are just scratching the surface of what's possible with collective intelligence.” ^[10]

Accuracy

The initial Google paper stated that the Google Flu Trends predictions were 97% accurate comparing with CDC data.^[4] However subsequent reports asserted that Google Flu Trends' predictions have been very inaccurate, especially in two high-profile cases. Google Flu Trends failed to predict the 2009 spring pandemic^[12] and over the interval 2011–2013 it consistently overestimated relative flu incidence,^[6] predicting twice as many doctors' visits over one interval in the 2012-2012 flu season as the CDC recorded.^[6]^[13] A 2022 study published (with commentaries) in the International Journal of Forecasting^[14] found that Google Flu Trends was outperformed by the recency heuristic, an instance of so-called "naive" forecasting, where the predicted flu incidence equals the most recently observed flu incidence. For all weeks from March 18, 2007, to August 9, 2015 (the horizon for which Google Flu Trends predictions are available), the mean absolute error of Google Flu Trends was 0.38 and of the recency heuristic 0.20 (both in percentage points; linear regression with a single predictor, the most recently observed flu incidence, had a mean absolute error of also 0.20, and the benchmark of random prediction had 1.80).

One source of problems is that people making flu-related Google searches may know very little about how to diagnose flu; searches for flu or flu symptoms may well be researching disease symptoms that are similar to flu, but are not actually flu.^[15] Furthermore, analysis of search terms reportedly tracked by Google, such as "fever" and "cough", as well as effects of changes in their search algorithm over time, have raised concerns about the meaning of its predictions.^[6] In fall 2013, Google began attempting to compensate for increases in searches due to prominence of flu in the news, which was found to have previously skewed results.^[16] However, one analysis concluded that "by combining GFT and lagged CDC data, as well as dynamically recalibrating GFT, we can substantially improve on the performance of GFT or the CDC alone."^[6] A later study also demonstrates that Google search data can indeed be used to improve estimates, reducing the errors seen in a model using CDC data alone by up to 52.7 per cent.^[17]

By re-assessing the original GFT model, researchers uncovered that the model was aggregating queries about different health conditions, something that could lead to an over-prediction of ILI rates; in the same work, a series of more advanced linear and nonlinear better-performing approaches to ILI modelling have been proposed.^[18]

However, followup work was able to substantially improve the accuracy of GFT through the use of a random forest regression model trained on both the incidence of influenza-like illness and the output of the original GFT model.^[19]

Related systems

Similar projects such as the flu-prediction project^[20] by the Institute of Cognitive Science at Universitat Osnabrück carry the basic idea forward, by combining social media data e.g. Twitter with CDC data, and structural models that infer the spatial and temporal spreading ^[21] of the disease.

References

External links

Google

Company

Divisions

Ads
AI
- Brain
- DeepMind
Android
China
- Goojje
Chrome
Cloud
Glass
Google.org
Health
Maps
Pixel
Search
- Timeline
Sidewalk Labs
Sustainability
YouTube
- History
- "Me at the zoo"
- Social impact
- YouTuber

People

Current	Krishna Bharat Vint Cerf Jeff Dean John Doerr Sanjay Ghemawat Al Gore John L. Hennessy Urs Hölzle Salar Kamangar Ray Kurzweil Ann Mather Alan Mulally Rick Osterloh Sundar Pichai (CEO) Ruth Porat (CFO) Rajen Sheth Hal Varian Susan Wojcicki Neal Mohan
Former	Andy Bechtolsheim Sergey Brin (Founder) David Cheriton Matt Cutts David Drummond Alan Eustace Timnit Gebru Omid Kordestani Paul Otellini Larry Page (Founder) Patrick Pichette Eric Schmidt Ram Shriram Amit Singhal Shirley M. Tilghman Rachel Whetstone

Real estate

Design

Fonts
- Croscore
- Noto
- Product Sans
- Roboto
Logo
- Doodle
  - Doodle Champion Island Games
  - Magic Cat Academy
Material Design

Events

Android Developer Challenge Developer Day Developer Lab Code-in Code Jam Developer Day Developers Live Doodle4Google G-Day I/O Jigsaw Living Stories Lunar XPRIZE Mapathon Science Fair Summer of Code Talks at Google
YouTube	Awards CNN/YouTube presidential debates Comedy Week Live Music Awards Space Lab Symphony Orchestra

Projects and
initiatives

20% project
Area 120
- Reply
- Tables
ATAP
Business Groups
Computing University Initiative
Data Liberation Front
Data Transfer Project
Developer Expert
Digital Garage
Digital News Initiative
Digital Unlocked
Dragonfly
Founders' Award
Free Zone
Get Your Business Online
Google for Education
Google for Startups
Labs
Liquid Galaxy
Made with Code
Māori
ML FairnessNative Client
News Lab
Nightingale
OKR
PowerMeter
Privacy Sandbox
Quantum Artificial Intelligence Lab
RechargeIT
Shield
Silicon Initiative
Solve for X
Starline
Student Ambassador Program
Submarine communications cables
- Dunant
- Grace Hopper
Sunroof
YouTube
- Creator Awards
- Next Lab and Audience Development Group
- Original Channel Initiative
Zero

Criticism

2018 data breach 2018 walkouts Alphabet Workers Union Censorship DeGoogle "Did Google Manipulate Search for Hillary?" Dragonfly FairSearch "Ideological Echo Chamber" memo Litigation Privacy concerns Street View San Francisco tech bus protests Services outages Smartphone patent wars Worker organization
YouTube	Back advertisement controversy Censorship Copyright issues Copyright strike Elsagate Fantastic Adventures scandal Headquarters shooting Kohistan video case Reactions to Innocence of Muslims Slovenian government incident

Development

Operating systems

Android
- Automotive
- Glass OS
- Go
- gLinux
- Goobuntu
- Things
- TV
- Wear OS
ChromeOS
- ChromiumOS
- Neverware
Fuchsia
TV

Libraries/
frameworks

Platforms

App Engine AppJet Apps Script Cloud Platform Anvato Firebase Cloud Messaging Crashlytics Global IP Solutions Internet Low Bitrate Codec Internet Speech Audio Codec Gridcentric, Inc. ITA Software Kubernetes LevelDB Neatx Project IDX SageTV
Apigee	Bigtable Bitium Chronicle VirusTotal Compute Engine Connect Dataflow Datastore Kaggle Looker Mandiant Messaging Orbitera Shell Stackdriver Storage

Tools

Search algorithms

Others

BERT BigQuery Chrome Experiments Flutter Gemini Googlebot Keyhole Markup Language LaMDA Open Location Code PaLM Programming languages Caja Carbon Dart Go Sawzall Transformer Viewdle Webdriver Torso Web Server
File formats	AAB APK AV1 On2 Technologies VP3 VP6 VP8 libvpx VP9 WebM WebP WOFF2

Products

Entertainment

Currents (news app) Green Throttle Games Owlchemy Labs Oyster PaperofRecord.com Podcasts Quick, Draw! Santa Tracker Songza Stadia games Typhoon Studios TV Vevo Video
Play	Books Games most downloaded apps Music Newsstand Pass Services
YouTube	BandPage BrandConnect Content ID Instant Kids Music Official channel Preferred Premium original programming YouTube Rewind RightsFlow Shorts Studio TV

Communication

Search

Aardvark
Alerts
Answers
Base
BeatThatQuote.com
Blog Search
Books
- Ngram Viewer
Code Search
Data Commons
Dataset Search
Dictionary
Directory
Fast Flip
Flu Trends
Finance
Goggles
Google.by
Images
- Image Labeler
- Image Swirl
Kaltix
Knowledge Graph
- Freebase
- Metaweb
Like.com
News
- Archive
- Weather
Patents
People Cards
Personalized Search
Public Data Explorer
Questions and Answers
SafeSearch
Scholar
Searchwiki
Shopping
Catalogs
- Express
Squared
Tenor
Travel
- Flights
Trends
- Insights for Search
Voice Search
WDYL

Navigation

Earth
Endoxon
ImageAmerica
Maps
- Latitude
- Map Maker
- Navigation
- Pin
- Street View
  - Coverage
  - Trusted
Waze

Business
and finance

Ad Manager
AdMob
Ads
Adscape
AdSense
Attribution
BebaPay
Checkout
Contributor
DoubleClick
- Affiliate Network
- Invite Media
Marketing Platform
- Analytics
- Looker Studio
- Urchin
Pay (mobile app)
- Wallet
- Pay (payment method)
- Send
- Tez
PostRank
Primer
Softcard
Wildfire Interactive
Widevine

Organization
and productivity

Bookmarks Browser Sync Calendar Cloud Search Desktop Drive Etherpad fflick Files iGoogle Jamboard Notebook One Photos Quickoffice Quick Search Box Surveys Sync Tasks Toolbar
Docs Editors	Docs Drawings Forms Fusion Tables Keep Sheets Slides Sites Vids
Publishing	Apture Blogger Pyra Labs Domains FeedBurner One Pass Page Creator Sites Web Designer

Education

Others

Account Dashboard Takeout Android Auto Android Beam Arts & Culture Assistant Authenticator Body BufferBox Building Maker BumpTop Cast Cloud Print Crowdsource Digital Wellbeing Expeditions Family Link Find My Device Fit Google Fonts Gboard Gemini Gesture Search Impermium Knol Lively Live Transcribe MyTracks Nearby Share Now Offers Opinion Rewards Person Finder Poly Question Hub Quick Share Reader Safe Browsing Sidewiki SlickLogin Sound Amplifier Speech Services Station Store TalkBack Tilt Brush URL Shortener Voice Access Wavii Web Light WiFi
Chrome	Apps Chromium Dinosaur Game GreenBorder Remote Desktop Web Store V8
Images and photography	Camera Lens Snapseed Nik Software Panoramio Photos Picasa Web Albums Picnik

Hardware

Smartphones	Android Dev Phone Android One Nexus Nexus One S Galaxy Nexus 4 5 6 5X 6P Comparison Pixel Pixel 2 3 3a 4 4a 5 5a 6 6a 7 7a Fold 8 8a Comparison Play Edition Project Ara
Laptops and tablets	Chromebook Nexus 7 (2012) 7 (2013) 10 9 Comparison Pixel Chromebook Pixel Pixelbook Pixelbook Go C Slate Tablet
Wearables	Fitbit List of products Pixel Buds Pixel Watch Pixel Watch 2 Project Iris (unreleased) Virtual reality Cardboard Contact Lens Daydream Glass
Others	Chromebit Chromebox Clips Digital media players Chromecast Nexus Player Nexus Q Dropcam Liquid Galaxy Nest Smart Speakers Thermostat Wifi OnHub Pixel Visual Core Search Appliance Sycamore processor Tensor Tensor Processing Unit Titan Security Key

v t e Litigation
Advertising	Feldman v. Google, Inc. (2007) Rescuecom Corp. v. Google Inc. (2009) Goddard v. Google, Inc. (2009) Rosetta Stone Ltd. v. Google, Inc. (2012) Google, Inc. v. American Blind & Wallpaper Factory, Inc. (2017) Jedi Blue
Antitrust	European Union (2010–present) United States v. Adobe Systems, Inc., Apple Inc., Google Inc., Intel Corporation, Intuit, Inc., and Pixar (2011) Umar Javeed, Sukarma Thapar, Aaqib Javeed vs. Google LLC and Ors. (2019) United States v. Google LLC (2020) United States v. Google LLC (2023)
Intellectual property	Perfect 10, Inc. v. Amazon.com, Inc. and A9.com Inc. and Google Inc. (2007) Viacom International Inc. v. YouTube, Inc. (2010) Lenz v. Universal Music Corp.(2015) Authors Guild, Inc. v. Google, Inc. (2015) Field v. Google, Inc. (2016) Google LLC v. Oracle America, Inc. (2021) Smartphone patent wars
Privacy	Rocky Mountain Bank v. Google, Inc. (2009) Hibnick v. Google, Inc. (2010) United States v. Google Inc. (2012) Judgement of the German Federal Court of Justice on Google's autocomplete function (2013) Joffe v. Google, Inc. (2013) Mosley v SARL Google (2013) Google Spain v AEPD and Mario Costeja González (2014) Frank v. Gaos (2019)
Other	Garcia v. Google, Inc. (2015) Google LLC v Defteros (2020) Epic Games v. Google (2021) Gonzalez v. Google LLC (2022)
Category

Terms and phrases	"Don't be evil" Gayglers Google (verb) Google bombing 2004 U.S. presidential election Google effect Googlefight Google hacking Googleshare Google tax Googlewhack Googlization "Illegal flower tribute" Rooting Search engine manipulation effect Sitelink Site reliability engineering YouTube poop
Documentaries	AlphaGo Google: Behind the Screen Google Maps Road Trip Google and the World Brain The Creepy Line
Books	Google Hacks The Google Story Google Volume One Googled: The End of the World as We Know It How Google Works I'm Feeling Lucky In the Plex The Google Book The MANIAC
Popular culture	Google Feud Google Me (film) "Google Me" (Kim Zolciak song) "Google Me" (Teyana Taylor song) Is Google Making Us Stupid? Proceratium google Matt Nathanson: Live at Google The Billion Dollar Code The Internship Where on Google Earth is Carmen Sandiego?
Others	"Attention Is All You Need" elgooG Predictions of the end Registry .app (top-level domain) .dev g.co .google Pimp My Search Relationship with Wikipedia Sensorvault Stanford Digital Library Project

Italics indicate discontinued products or services.
Category
Commons
Outline
WikiProject