← Back to all posts

On Data Value: the only V that justifies the platform

Value is the only V that justifies the existence of the data platform. Data has value when it informs a decision that produces a better outcome. Most data has no value. It was collected because it could be, stored because storage is cheap, and never used. Useless data is a liability, not an asset.

data-engineeringvalueeconomicsroiprioritization

Value is not a property of data. It is a property of the decision the data informs. Data that informs no decision has no value, regardless of how expensive it was to collect, how sophisticated the pipeline that serves it, or how impressive the dashboard that displays it.

Value is the only V that justifies the data platform's existence. Data has value when it changes a decision. The value is the difference between the outcome with the data and the outcome without it. If the data doesn't change the decision, it has no value. If the decision would have been the same either way, the data added nothing. The nothing was expensive to produce.

An e-commerce company collects every click, page view, add-to-cart, and purchase. The clickstream alone is 10 TB per day. The data team builds pipelines to ingest it, transform it, serve it. Six months later, nobody queries the clickstream tables. The analysts query purchases and ignore the rest. The clickstream data has zero value. The storage costs $2,400 per month. The pipeline maintenance costs one engineer at 20% time. The total cost exceeds the value. The data is a liability.

Most data in most organizations is a liability. It was collected because it could be — the event was there, the tracking was easy, the storage was cheap. It was never used because nobody knew what question it answered. The question came first: "we should track everything, we might need it later." The question was wrong. The right question is: "what decision will this data inform, and what is the value of a better decision?" If the answer is unclear, don't collect the data. The collection is not free. The cost is the pipeline, the storage, the maintenance, the cognitive load on the data team, the dilution of the data catalog with tables nobody queries. The cost is paid forever. The value is zero. The net is negative.

Value forces prioritization. You cannot pipeline every data source. You must pipeline the sources that will answer the business's most important questions. Identifying those questions requires talking to the people who will use the data. The talking is the most underinvested activity in data engineering. Engineers build pipelines for data they have, not for questions the business needs answered. The pipeline exists. The question doesn't. The value is zero.

Doug Laney's Infonomics (2017) proposes treating data as a balance-sheet asset: measure its value, depreciate it over its useful life, account for the cost of maintaining it. The proposal is radical because almost no organization does it. Data is treated as free. It is not free. The cost is real. The value is real only when the data is used. The gap between cost and value is the economic problem of data. Closing the gap is the discipline of data economics.

See: Doug Laney, "Infonomics: How to Monetize, Manage, and Measure Information as an Asset for Competitive Advantage" (Routledge, 2017). Thomas C. Redman, "Data's Credibility Problem" (Harvard Business Review, 2013). Benn Stancil, "The Data Platform Cost Model" (Mode, 2020).

This post is part of a series on The Many Vs of Data, originating from Doug Laney's 2001 Gartner note. Each V names a dimension of why data is hard.