Guides

Housing and community data, explained

How housing markets and neighborhood conditions get measured — the indexes, the record-level data behind them and who publishes what for free.

Housing sits at the intersection of two very different kinds of data. One side is transactional and financial: what did this specific property sell for, what is it worth now, who owns it. The other is social and geographic: who lives in this neighborhood, can they afford to stay, is the community gaining or losing population. Analysts working on housing policy, community development or real estate almost always need both, pulled from different sources and built on different units of geography.

The questions people ask

A city planner asks which neighborhoods are at risk of displacement as prices rise. A lender or appraiser asks what a specific property is worth today. A nonprofit or foundation asks where housing cost burden is concentrated, to target assistance. A journalist or economist asks whether the national housing market is cooling. Each of these pulls from a different layer of the data stack — property records, market indexes, or census demographics — and conflating the layers is a common source of confused analysis.

The data it runs on

Three distinct data types make up this field. Market-trend data — home value indexes, rent indexes, inventory and days-on-market — is aggregate and geographic, published at the national, state, metro, county or ZIP level, and answers "what is happening to prices in this area," not "what is this specific house worth." Property-level records — deeds, tax assessments, mortgages, foreclosures — are transactional and parcel-specific, sourced from county and municipal recorders and licensed in bulk by data vendors rather than browsed one property at a time. Demographic and community data — population, income, housing cost burden, tenure — comes primarily from the American Community Survey and the decennial census, aggregated to geographies like the census tract, a small, relatively stable statistical subdivision designed to approximate a neighborhood.

Core metrics and how to read them

A house price index tracks how home values change over time for a consistent set of properties, typically by comparing repeat sales or appraisals of the same homes rather than simply averaging all sale prices in a period — a distinction that matters because a raw average sale price shifts whenever the mix of homes selling changes (more starter homes selling one month, more luxury homes the next), which a repeat-sales index is specifically designed to avoid.

The housing affordability index measures whether a typical household's income is sufficient to qualify for a mortgage on a typical home at current prices and interest rates, combining three moving parts — income, home prices, and financing cost — into one number, which means the same affordability reading can result from very different underlying conditions (high prices and low rates, or lower prices and high rates).

Rent burden measures the share of a household's income spent on rent, with 30% as the traditional (if debated) threshold for "cost-burdened." It is one of the most widely cited housing-hardship metrics precisely because it's calculated from a single, consistently collected data source — the American Community Survey — down to the tract level.

Vacancy rate measures the share of housing units unoccupied at a point in time, and its interpretation depends heavily on context: a high vacancy rate can signal oversupply and weak demand in one market, or, in a market with heavy seasonal or second-home ownership, simply reflect normal seasonal patterns rather than distress.

An automated valuation model (AVM) estimates a specific property's value using statistical models trained on comparable sales, tax assessments and property characteristics, without a human appraiser physically inspecting the property. AVMs are fast and scalable — the mechanism behind instant online home-value estimates — but they're a model output, not an appraisal, and their accuracy varies with how many comparable, recent sales exist nearby; they tend to be least reliable in thin markets or on atypical properties.

Displacement risk combines several of the above — rising prices, rent burden, demographic change, proximity to new investment — into an index intended to flag neighborhoods where existing residents are at elevated risk of being priced out. It's inherently a composite, judgment-laden metric rather than a single clean observation, and different organizations building "displacement risk" indexes make different, defensible choices about which inputs to weight most heavily.

How the work is done in practice

For free, aggregate market-trend data, Zillow Research and Redfin Data Center are the two standard references economists and journalists reach for first: both publish downloadable home value, rent, inventory and price-trend datasets by geography at no cost, drawing on each company's own listings data, and differ mainly in methodology and update cadence rather than in being fundamentally different products.

For property-level and valuation work, ATTOM and HouseCanary serve different parts of the same problem. ATTOM is a data licensing business rather than an end-user analytics tool: it aggregates nationwide tax, deed, mortgage and foreclosure records into a standardized model that other companies build products on top of. HouseCanary is narrower and more direct: automated valuation models, rental estimates and forecasts delivered through self-serve reports or a usage-priced API, aimed at agents, appraisers, lenders and investors who want a valuation on demand rather than a bulk licensing relationship.

For demographic and community data, U.S. Census Bureau Data Tools is the authoritative free source — data.census.gov, the public Census API, and geographic tools like TIGERweb — covering the decennial census and the American Community Survey down to small geographies at no cost. PolicyMap builds on top of that same public data (plus thousands of other licensed datasets), pre-cleaned and mapped down to the neighborhood or tract level, aimed at nonprofits, foundations and government agencies who want to skip the data-wrangling step census data otherwise requires.

Common mistakes and misreadings

Reading a raw median or average sale price as a price index. Without controlling for which homes sold in a given period, a rising "average price" can simply mean more expensive homes changed hands, not that home values broadly rose — the reason repeat-sales indexes exist.

Treating an AVM estimate as an appraisal. An automated valuation model is a statistical estimate with an error range that widens in thin or unusual markets; lenders and courts generally still require a licensed appraisal for anything consequential.

Comparing vacancy rates across very different markets without context. A resort town's vacancy rate reflects seasonal ownership patterns, not housing distress, and reading it the same way as a year-round urban market's vacancy rate produces a false comparison.

Using census-tract data as if tracts were neighborhoods with fixed boundaries. Tracts are a statistical convenience redrawn periodically by the Census Bureau, and their boundaries don't always match how residents actually define their own neighborhood.

Building or citing a displacement-risk score without checking its inputs. Because displacement risk indexes are composite and judgment-laden, two organizations' indexes for the same city can disagree meaningfully, and neither should be treated as a single objective ground truth.

For the full landscape of tools covering property data, demographics and urban planning, see every real estate analytics tool in this category and every demographic data tool in this category.

Related tools

Terms used in this guide

Latest on this topic