How we calculate our market data
Every number on imotesa.bg comes from public listings put through the same process. This page describes that process — including what the data cannot show.
1. Data sources
The data comes from publicly available Bulgarian property listings published on the major property portals. We do not use Registry Agency records, notarial deeds or any other register of completed transactions — see “Limitations”.
Crawling is automated and repeats every hour. A listing not seen for 24 hours is marked inactive and leaves every statistic. So a listing pulled from its portal last night may still count as active for part of today.
Active listings are the only basis for the market figures. Inactive ones are kept for price history but take no part in the medians.
Current coverage
- 210,000+Active Listings
- 14,000+Merged Duplicate Listings
- 440+Neighbourhoods with AI Score
- 250+Cities & Regions
2. Duplicate detection
The same apartment is often published by several agencies across several portals at once. Counted as separate properties, they skew every statistic towards whatever is advertised most aggressively.
Listings are therefore grouped into a single canonical property on two independent signals:
- Perceptual image hashing — each photo is reduced to a short fingerprint that survives resizing, cropping and watermarking. Two listings showing the same home are recognised even when the image files differ.
- Attribute matching — location, floor area, storey, year built and price. An image match without an attribute match is not enough — that is what stops a developer's stock photography from merging different apartments in one building.
The merged property keeps all of its listings, so its page shows which agency offers it at which price. Where the prices differ, one listing per property feeds the market calculations — not the average of the duplicate postings.
Detection is not infallible. A listing with no photos and sparse attributes may survive as a separate property; the opposite error — merging two near-identical neighbouring apartments — is rarer but possible.
3. Median, not mean
The published figures are medians: the value with an equal number of listings above and below it. A mean is dragged upward by a small number of very expensive properties and would describe a market most buyers would not recognise. The median describes the typical listing.
Page headings say “average prices” because that is how people ask the question. The figures themselves are labelled as medians everywhere a specific value is asserted.
A segment is published only when it holds enough listings for the median to be stable. The thresholds differ by property type:
| Property type | Minimum listings |
|---|---|
| Apartments, studios, single-room units | 15 |
| Houses and house floors | 8 |
| Parking spaces and garages | 5 |
A segment below its threshold is not published at all — no page, no number, no estimate. That is why some neighbourhoods have apartment statistics but no house statistics.
Segments are calculated at four levels: neighbourhood and room count, neighbourhood overall, city and room count, city overall. Each level must clear the threshold on its own.
4. The “% below market” signal
The badge on a listing compares its price per square metre against the median price per square metre for the segment the property falls into:
deviation = (segment median €/m² − listing €/m²) ÷ segment median €/m²
The most specific segment with sufficient data is used, in this order: neighbourhood and room count → neighbourhood (all rooms) → the city as a whole. That last step is why properties in towns with no named neighbourhoods, and types with no room count (houses, parking), still receive a signal.
The badge's thresholds are deliberately asymmetric: below −8% the property reads as cheaper than the market, above +10% as more expensive, and a gap of more than 40% below market is shown as a warning rather than a better deal. A gap that large usually means a data error or a defect the listing does not mention.
This is not a valuation. The signal accounts for none of condition, aspect, storey, noise, view, build quality or legal status — the things that explain much of the price difference. It shows where the asking price sits relative to neighbouring listings, and nothing more.
5. Limitations
Each of these is inherent to the source rather than a defect awaiting a fix:
- Asking prices, not transaction prices — we measure what sellers want, not what deals close at. When the market cools, asking prices lag actual ones by months.
- Coverage is not the whole market — properties sold without a listing — between acquaintances, through a single broker, or before publication — take no part. In new developments a share of sales never reaches a portal at all.
- AI enrichment can be wrong — property type, attributes and summaries are generated automatically from the listing's text and photos. Where the listing is inaccurate or incomplete, the error carries through.
- Period comparisons are sensitive to mix — a quarter-on-quarter change reflects which properties are currently on offer, not only a change in the price of the same property. We therefore show a trend only when the sample is sufficient in both periods.
6. Citation and reuse
The data may be freely cited — by journalists, analysts, students and researchers — with attribution and a link to the specific page the figure came from.
Please do not present our figures as completed-transaction prices, and please state the date they were valid on: the market pages are recalculated daily.
Example citation
imotesa.bg, “Average prices for apartments for sale in Sofia”, data as of 10 August 2026, https://www.imotesa.bg/statistics/sale/apartment/sofiya
Questions about the methodology, or a breakdown that is not published: support@imotesa.com
Methodology last revised: August 2026