Science, Data, and Knowledge: End of Theory and Understanding the World

Rianna Herzlinger

Volume 1 • Issue 1

The popularity and reliance on big data has affected many areas of technology and knowledge. One of the biggest changes has occurred in the scientific disciplines, with a shift away from scientific theory and towards data analytics. ‘End of Theory’ describes and supports this shift in how scientific knowledge is accumulated. This paper provides an overview of big data, the rise of End of Theory through case studies, and potential solutions to integrate theory and data. It ultimately concludes that even in ideal circumstances, sacrificing theoretical knowledge for data-driven predictions hinders us from understanding the world we live in.

1) Big Data Creates Big Hype

The term ‘big data’ refers to the growing amounts of data that can be gathered and analyzed through our use of the internet and new technology. With its increased ‘volume, velocity, and variety,’ big data posed challenges to data scientists, being too large and unwieldy to be assessed via traditional statistical methods (Laney, 2001). The ‘problem of big data’ most broadly refers to the challenges of managing and analyzing complex and massive datasets. However, since the advent of deep learning (DL), the problems of big data have shifted. Deep learning refers to larger and computationally complex neural networks, which can process massive amounts of raw, numerical data and identify patterns. Once a DL model is trained on data and can recognize statistical patterns in the data, it can use this information to make predictions about unseen data. Deep learning does better with more data, meaning that the success of deep learning— or artificial intelligence— is reliant on big data itself.

Now that big data can be used, the problems are centered around data quality and origin. For example, researchers have pointed to the harm of applying statistical techniques to data that has been repurposed or consolidated from multiple sources (Clarke, 2016). The fear is that without proper knowledge of the origin and context of data, statistical models risk misrepresenting dependencies between data. Others worry about ‘passive data collection’ — a tendency to cast a broad net and acquire all the data possible without active understanding of the relationship between data (Fricke, 2015). Biased data results from unrepresentative data collection or information incompleteness/noise (Liu et al., 2016). If data isn’t collected without understanding what features of the studied population are important, samples may fail to be representative of the population at issue. Thus, any patterns discovered by the model may be noise rather than indicative of the true population, resulting in consistency and reliability issues (Coveney et al., 2016). Unreliable and inconsistent data creates ethical issues when discoveries by these models are used to make policy or practical decisions.

Socio-political fears about big data include privacy concerns and human resource scarcity (Hilbert, 2016; Boyd and Crawford, 2012). Some fear that data collection could be used to track acts of civil disobedience and suppress free speech (Boyd and Crawford, 2012). There is further concern about the lack of equity between those who collect, store, and utilize big data versus those whom data collection targets (Andrejevic, 2014). A prime example of this fear is the predatory nature of social media, which siphons data from its users to sell it to third party companies. Data forcibly collected from users without any compensation is then used to support invasive marketing techniques as these third party companies create targeted advertisements for users (Boyd and Crawford, 2012). Thus, even if big data creates good quality predictions, its applications create other social and ethical concerns.

While big data has affected many fields of knowledge and technology, it has had a unique impact on scientific epistemology— the ways that we create and validate knowledge in scientific domains (Keller, 2022). Knowledge in the physical sciences has been traditionally fostered via the scientific method. Attributed to Francis Bacon in the 17th century, modern science is a fusion of observation and reason (Coveney et al., 2016). The scientific method proceeds as follows: researchers identify research gap, create a research question from existing or extended theory, formulate hypotheses to address their question, design studies to minimize confounding variables, collect data, and analyze data to draw inferences (Coveney et al., 2016; Maass et al., 2018). Data driven research, on the other hand, begins with a research question but immediately proceeds by creating/obtaining sources of data, cleansing and extracting data streams, integrating, aggregating, and representing data to detect insights, analyzing and modeling data, and interpreting these patterns (Maass et al., 2018). These methods differ both in their grounding and their processes. Theory-driven research relies on previous theories about the discipline; while data-driven research can be similarly based, it can easily skip this grounding and instead move directly to data collection. Second, data-driven research isn’t particularly concerned with why or how patterns result in correct predictions, simply that they do create correct predictions.

2) Big Data at Its Best: End of Theory

Despite concerns over the misuse of big data, some have embraced big data wholeheartedly. Support of big data led to a save of ‘End of Theory’ enthusiasts. The term “End of Theory” first emerged in an explosive WIRED article, published by Chris Anderson in 2008. In this article, Anderson describes us as having entered the Petabyte Age: an era in which we have access to massive amounts of data. As a result, the companies that leverage this data no longer know why phenomena occur— it is enough to know they occur. This extends to predicting human reactions to advertisements, creating translation systems, and matching ads to content. Humans have become more predictable, but arguably not more understandable. This attitude has extended to scientific innovation too. Where the traditional science involved a hypothesis, model, and testing, this process has become obsolete. As Anderson writes, “Petabytes allow us to say: ‘Correlation is enough.’” Thus, the end of theory refers to an era in which scientific hypothesizing and modeling is subordinated to the collection and analysis of massive quantities of data.

Anderson himself appears enthusiastic about our future without theory. He ends his article with the chilling statement: “There’s no reason to cling to our old ways. It’s time to ask: What can science learn from Google?” Part of his criticism of ‘the old ways’ is that the scientific models we relied on were really simplifications of phenomena, rarely correct but still somewhat useful. These models must be constantly refined when scientists encounter new facts about the world that violate their models. Anderson sees this as a cumbersome and unhelpful process— especially when the alternative is easier and often more accurate. Part of the explosive nature of this article is not Anderson’s identification of this trend, but his want to fully embrace a post-theory world.

What is the value of theoretical knowledge? Anderson makes a valid point about traditional scientific models: they aren’t strictly true. They are necessary simplifications of the real world. But does that make them invaluable? I argue theoretical knowledge is still incredibly valuable. First, their simplicity allows them to be applied in a wide variety of cases. Theoretical knowledge allows us to extrapolate to unseen cases and have scientific discovery that builds on previous discoveries. These are attributes that are missing in data-driven research. Karpatne et al. (2017) comment that there are two aspects of scientific knowledge discovery that make it difficult for data-driven models to be highly effective. First, in many of the physical sciences, the variables within a system are complex and dynamic. With limited labeled data, it becomes difficult to represent the true nature of the relationships between variables in scientific problems. This leads to difficulty in generalizing to wider cases. Others echo the inability of these DL models to extrapolate beyond existing data (Coveney et al., 2016). Because these models track observable patterns, they aren’t designed to model structural characteristics of the underlying system. Without this theoretical or foundational knowledge, they struggle to deal with examples beyond the range of their training data. Thus, theoretical knowledge is crucial to creating foundational and widely-applicable understanding of a phenomenon.

However, the biggest problem with data-driven methods is the difficulty of extracting theoretical knowledge from their results. Often termed ‘black boxes’, DL models do not understand its input data as real-world phenomena but rather strings of numbers. Karpatne et al. (2017) argue that this inability stems from the disparate goals of DL models vs theory. The end goal of a DL model is the generation of actional models, predictions that match reality. However, science doesn’t end there: science takes these predictions and creates interpretable theories from them. Scientific theory is interested in questions that data-driven science doesn’t ask. Simply put, theory wants to know why. However, DL models cannot help with these questions.

These criticisms take End of Theory on its own terms: they assume the mutual exclusivity of data-driven and theory-driving discovery. However, other researchers criticize one of the assumptions made in Anderson’s argument: that big data doesn’t need theory to begin with. But data science— good data science— requires theoretical knowledge to make their models in the first place. Why collect some information and not others? Why employ this learning algorithm and not that one (Pigliucci, 2009)? A theoretical understanding of the present problem must guide data collection, curation, and interpretation (Coveney et al., 2016). Pigliucci (2009) concludes: “Without models, mathematical or conceptual, data are just noise,” (1). Thus, Anderson’s support of End of Theory rests on a fallacy: that big data can actually create good results without first having a theoretical understanding of the problem.

Coveney et al. (2016) gives a potent example of how medicine requires theory. Take a DL model which predicts a patient’s reaction to a drug. It predicts the outcome of an individual by comparing it with the outcomes of an aggregated population. While it could be right most or all of the time, the model itself does not help doctors understand why each patient is likely to be helped by a specific drug. Thus, reliance on a DL model to combat this problem actively debilitates the creation of personalized or precision medicine. If a patient is outside the general distribution of reactions or is from an underrepresented group, the DL model will likely mispredict their reactions to drugs. In a fully End of Theory world, doctors will lack the know-how to help patients when data fails; they won’t even know why the failure happened. Without understanding the causality behind health decisions, doctors cannot create treatments tailored to a patient’s biology.

Even when big data does everything right— when DL model predictions are correct 100% of the time, data is wholly representative, and data is sourced strategically— it still fails to produce understanding of the world.

3) Case Studies

The following two case studies illustrate the value of theory, state of theory in DL models, and the consequences of a reduced understanding of phenomena in the world.

a) Google Flu Trend

In 2013, Google Flu Trend (GFT) became famous for seemingly predicting the 2012-13 flu season better than the CDC. They used data from users’ search patterns. They used an unsupervised model tasked with making the best matches among 50 million searches to fit 1152 clusters the model itself had to discover (Poeter, 2014). However, aspects of the search engine itself diluted the outcomes. For example, the autosuggest function in Google led to altered user searches that the model didn’t account for. Further, it overfit to seasonal terms, picking up on unrelated terms like ‘high school basketball’ as related to the flu because it jumped in searches during a particular season. Although GFT has since updated their model, the damage of the first GFT was already done.

However, after the hype died down, it became clear that GFT actually overpredicted flu cases by 50% each year (Lazer et al., 2014). The spectacular claim that devolved into error led academics to coin the term “big data hubris” to describe Google’s failure. What this article describes as data hubris is essentially an implicit belief in the end of theory. GFT used no theoretical knowledge to sort its clusters or identify search inquiries with the flu. Instead, it used the brute force of massive amounts of data and a black box algorithm to categorize these searches as indicative/not indicative of the flu. The article critiques that big data is often seen as “a substitute for, rather than a supplement to, traditional data collection and analysis” and that “we are far from a place where we can supplant more traditional methods or theories,” (page 1).

With these criticisms, let us return to the thematic questions. First, what is the value of theoretical knowledge? This example further highlights this value because the bar is so low. For example, the theoretical knowledge at stake here is simply the understanding that ‘high school basketball’ searches do not bear on flu rates. We are far from the realm of protein sequencing or physics modeling. However, this elementary conclusion is missed by a model that processes data without understanding its context or sentiment. No theory can be generated from this model, even if it correctly predicted flu rates. It has no way to discern which of its search queries resulted in someone that matched the description of the flu and further in someone who actually had the flu. At best, it diagnoses at one step further away from symptomatology by focusing not on the symptoms themselves but an inquiry into the symptoms. This does not create the type of causality required for a good theory. Finally, the consequences of mistakes like GFT are stark and grave. GFT was proven inaccurate because of comparison to reputable and carefully collected data from the CDC. We had not transitioned into making decisions based on the GFT results, thus there were no severe negative consequences to this error. However, in a true post-theory world, there would be no CDC to compare the GFT to. In this case, we risk making important, costly, and live-altering medical decisions based on an algorithm we do not understand and cannot trust. These consequences ought to give pause to advocates of the end of theory.

b) AlphaFold

The third iteration of Alpha Fold, first created in 2018, Alpha Fold 3 can now predict not only single-chain proteins but also the structure of protein complexes within DNA, RNA and more. Co-developed by Google DeepMind, it leverages the most recent ML architecture: transformers. The model begins with clouds of atoms and refines their positions to generate 3D representations of molecular structures. Before this technology, predicting the folding of a protein from its amino acid sequence was possible but a deeply time-consuming and painstaking process, taking an entire PhD to produce just one. Now, Alpha Fold 3 can do this instantly and with immense accuracy.

Alpha Fold is causing a tearing in scientific discipline, as the historical value of confirmation/validation is being subordinated to blind belief in Alpha Fold’s predictions. For example, seasoned academics report PhD students give presentations without even denoting that the structure they’re referencing is a prediction. Sometimes, Alpha Fold can make predictions that don’t match up to real activity changes between proteins. So what? Weren’t the original theoretical models for these proteins also inaccurate? Yes, but scientists understood why these models were true or untrue. With Alpha Fold, we cannot say why any protein is or isn’t predicted to fold a certain way. Thus, the validation process itself isn’t helped— scientists still need to do the painstaking work of proving the prediction. Worse, as a new generation of scientists work and contribute to their discipline post Alpha Fold, they may lose appreciation for validation entirely.

Alpha Fold demonstrates why theoretical knowledge is important. First, because it can help with the generation of further knowledge. Understanding how one protein folds creates a knowledge base with which to understand another protein. Second, because it contributes to genuine understanding of the knowledge we have. Scientists don’t just know how that protein folds as it does, they can explain why it does and under what conditions it would react otherwise. Alpha Fold does not contribute to this theoretical knowledge. The validation process of these structures still requires the same process as it did pre-ML. Currently, seasoned scientists treat these predictions as preliminary, unproven and suspect. Papers make statements like “a prediction can still provide a useful starting hypothesis, but it is even more important to seek independent experimental data to validate conclusions,” and that Alpha Fold is an “action plan for future development,” (Kovalevskiy et al. 2024). This recognizes the contributions ML models can make to science but caution they are not the ends of scientific knowledge, merely the means. But, in an end of theory world that Anderson advocates for, scientists would no longer see a reason to validate Alpha Fold’s results; it would become gospel. What would that mean for science? It leaves us with mistakes we have no understanding of how to fix. It leaves us with drugs that could have serious consequences that we cannot augment because we don’t know where the prediction went wrong. And worse, it leaves us with a fragile and fleeting understanding of the proteins themselves.

4) Big Solutions

RETURN TO THE ISSUE

The Simulated Turn

Volume 1 • Issue 1

If I’ve done my job well, you should be very concerned about the encroachment of data-driven research on the scientific community. Where do we go from here?

Academics have pushed back against End of Theory, advocating for ‘big theory’ to match the rise of big data (Boellstorff, 2013; Karpatne et al., 2017; Maass et al., 2018). While these advocacies come in many forms, they all agree that science must combine data-driven and theory-driven research. Maass et al. (2018) suggest an information systems framework to help move patterns to theory and theory to patterns. Assessing DL models with known theory can help explain why a model failed and how to improve it. More broadly, it can use model outcomes to create general theories that go beyond the limits of specific data. Starting with theory, domain theorists can suggest novel sources of data or combinations of data to facilitate data requirements.

Karpatne et al. (2017) gives many ideas for hybrid models: models that combine theory and data analytics. One possibility is the creation of physically consistent models. These are supervised DL models that are trained to be consistent with scientific principles. The authors suggest that physical consistency should be considered a critical aspect of model performance. Testing models for physical consistency can help prune models that can’t create consistent results. Within a model family, learning algorithms and optimization metrics should be considered for their physical consistency. This could take the form of initializing models with physical rules in pre-training, creating domain-guided constraints during the training process, or encoding scientific knowledge in probabilistic relationships. All of these techniques would require the combination of domain-specific expertise (the theory) and technical expertise (the data) to create hybrid models. Outputs can be further refined and validated with scientific knowledge.

In good news, science has already begun to adapt. Hybrid approaches are already being used in mapping climate change patterns, closing knowledge gaps in turbulence modeling efforts, discovery of novel compounds in material science, density functionals in quantum chemistry, better imaging tech in biomedical science, discovery of genetic biomarkers, estimation of surface water dynamics on global scale, (Karpatne et al., 2017).

While shifts in scientific epistemology may be concerning, theory is far from obsolete. Ultimately, science depends on explaining and understanding results. To pursue science without this impulse will inevitably erode the legitimacy of the discipline itself. But the world is full of both unknown things and people eager to know them. As these people continue to gravitate towards scientific research, they will be met with a changing discipline. But science need not abandon theory to succeed; it can leverage big data in existing theoretical frameworks to make new discoveries that continue to contribute to our knowledge of the world.

Works Cited

“Alphafold.” Wikipedia, Wikimedia Foundation, 19 Feb. 2025,
https://en.wikipedia.org/wiki/AlphaFold.

Anderson, Chris. “The End of Theory: The Data Deluge Makes the Scientific Method Obsolete.” Wired, Conde Nast, 23 June 2008, https://www.wired.com/2008/06/pb-theory/.

Andrejevic, Mark. “Big data, big questions| the big data divide.” International Journal of Communication 8 (2014): 17.

Boellstorff, Tom. “Making big data, in theory.” First Monday 18.10 (2013).

Boyd, Danah, and Kate Crawford. “Critical questions for big data: Provocations for a cultural, technological, and scholarly phenomenon.” Information, communication & society 15.5 (2012): 662-679.

Clarke, Roger. “Big data, big risks.” Information Systems Journal 26.1 (2016): 77-90.

Coveney, Peter V., Edward R. Dougherty, and Roger R. Highfield. “Big data need big theory too.” Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 374.2080 (2016): 20160153.

Frické, Martin. “Big data and its epistemology.” Journal of the association for information science and technology 66.4 (2015): 651-661.

Graham , Mark. “Big Data and the End of Theory?” The Guardian, Guardian News and Media, 9 Mar. 2012, https://www.theguardian.com/news/datablog/2012/mar/09/big-data-theory.

Hilbert, Martin. “Big data for development: A review of promises and challenges.” Development Policy Review 34.1 (2016): 135-174.

Jumper, John, et al. “Highly Accurate Protein Structure Prediction with Alphafold.” Nature News, Nature Publishing Group, 15 July 2021,
https://www.nature.com/articles/s41586-021-03819-2.

Karpatne, Anuj, et al. “Theory-guided data science: A new paradigm for scientific discovery from data.” IEEE Transactions on knowledge and data engineering 29.10 (2017): 2318-2331.

Keller, E. F. “Models, Simulation, and ‘Computer Experiments.’” In H. Radder
(Ed.), The Philosophy of Scientific Experimentation. University of Pittsburgh Press, 2022.

Kovalevskiy, Oleg, et al. “AlphaFold Two Years on: Validation and Impact.” PNAS, 12 Aug. 2024, https://www.pnas.org/doi/10.1073/pnas.2304819120.

Laney D. (2001) 3D data management: controlling data volume, velocity and variety. Meta-Group, February 2001, at http://blogs.gartner.com/doug-laney/files/2012/01/ad949-3D-Data-Management-Controlling-Data-Volume-Velocity-and-Variety.pdf.

Lazer, David, et al. “The Parable of Google Flu: Traps in Big Data Analysis.” Policy Forum, 14 Mar. 2014.

Liu, Jianzheng, et al. “Rethinking big data: A review on the data quality and usage issues.” ISPRS journal of photogrammetry and remote sensing 115 (2016): 134-142.

Maass, Wolfgang, et al. “Data-driven meets theory-driven research in the era of big data: Opportunities and challenges for information systems research.” Journal of the Association for Information Systems 19.12 (2018): 1.

Pigliucci, Massimo. “The end of theory in science?.” EMBO reports 10.6 (2009): 534-534.

Poeter, Damon. “Big Data Hubris: How Google’s Flu Tracker Went Wrong.” PCMAG, PCMag, 14 Mar. 2014, https://www.pcmag.com/news/big-data-hubris-how-googles-flu-tracker-went-wrong.

R/Labrats on Reddit: People Are Overestimating Alphafold and It’s a Problem, https://www.reddit.com/r/labrats/comments/1b1l68p/people_are_overestimating_alphafold_and_its_a/. Accessed 21 Feb. 2025.