A Scraped Frame: Coverage, Entity Resolution and Non-Random Disclosure in Job Posting Data
← Chapter 179
Capstone 19 · Technical Report
Technical Report

A Scraped Frame: Coverage, Entity Resolution and Non-Random Disclosure in Job Posting Data

A census of a frame nobody designed, benchmarked against an external source.

Author  John Fisher
Series  Statistics, Data Science and AI: A Visual Handbook
Design  Census of one public job board, one quarter, one region
Where this comes from
Chapter Chapter 179 · Web Scraping as a Sampling Frame: Job Postings
Part Part XXVIII · Capstone Projects: Sampling & Data Collection
Dataset capstone-web-scraping-job-postings.xlsx
Notebook View the analysis
Abstract

Objective. To estimate advertised salary and remote-working prevalence for data-analyst roles in one region, and to characterize the error structure of a scraped sampling frame. Methods. All listings matching the occupation on one public board were captured over one quarter under a documented permissions protocol, yielding 1,385 rows. Because collection was exhaustive within the frame, sampling error is nil and the analysis concerns coverage, duplication and item non-response. Listings were resolved to vacancies; salary was reweighted to the board's and then to the region's seniority distribution; results were benchmarked against an official occupational earnings survey. Results. 1,382 unique listings resolved to 808 vacancies, so listing counts overstate vacancy counts by 71.0%. Salary was disclosed for 63.1% of vacancies, with disclosure falling monotonically from 86.3% (Junior) to 23.4% (Lead) while mean disclosed pay rose from $50,171 to $101,787, establishing missingness not at random. The complete-case mean was $64,153 against a benchmark of $71,800 (-10.6%). Reweighting recovered 56% and then 73% of the gap. Remote-working prevalence was accurate to -100.0%. Conclusions. Exhaustive collection eliminates the only error component with a standard formula and leaves the others untouched. Estimand validity must be argued variable by variable.

Keywords: Web scraping; sampling frame; coverage error; entity resolution; missing not at random; complete-case analysis; post-stratification; external benchmarking.

1. Introduction

Datasets assembled by automated collection are now routine, and they are commonly treated as populations rather than as samples on the grounds that collection was exhaustive. The inference is valid as far as it goes: a census of a frame has no sampling error. It is also the least useful thing that can be said about such a dataset, because the frame was constructed by processes external to the research, and the remaining components of total survey error are unaffected by the completeness of collection.

This report treats a scraped corpus of job advertisements as what it is, a census of a frame with unknown selection properties, and quantifies each error component against an external benchmark. The exercise is possible only because such a benchmark exists; the more general point is that without one, none of the errors documented below would be detectable from the data.

2. Collection protocol

Table 1. Collection specification. The permissions record belongs in the method section, not in an appendix.
ElementSpecification
Target populationAll open data-analyst vacancies in the region during the quarter
FrameListings visible on one public board, not excluded by robots.txt, reachable within the rate limit
MethodExhaustive capture of matching listings. A census of the frame, not a sample
Rate limitOne request per two seconds; identifying user agent with contact address
AccessNo authentication walls circumvented
Personal dataNamed recruiter details discarded at parse time, not retained then removed

The distinction between what is technically achievable and what is permitted is not enforced by software. A collection protocol that cannot be stated should not be published, on the same principle that governs the reporting of ethical approval in human-subjects research.

3. Entity resolution

Table 2. Reduction from listings to vacancies.
StageRows
Raw capture1,385
After removing crawler duplicates1,382
Resolved to distinct vacancies808
Inflation if unresolved71.0%

A single vacancy may be advertised repeatedly by the employer and concurrently by one or more agencies operating under their own trading names, with divergent job titles. The unit of analysis is therefore not observable directly in the raw data and must be reconstructed. Reporting listing counts as vacancy counts overstates market volume by more than seventy percent, an error of magnitude far exceeding any sampling consideration.

4. Item non-response on the salary field

Table 3. Salary disclosure by seniority. Disclosure decreases monotonically as pay increases.
SeniorityVacanciesDisclosure rateMean disclosed salary
Junior21186.3%$50,171
Mid36864.4%$66,020
Senior16546.1%$84,388
Lead6423.4%$101,787

The disclosure mechanism depends on the value of the variable being disclosed, which is the defining condition of missingness not at random. Complete-case analysis under MNAR is biased, and the direction of the bias is predictable from the mechanism: the unobserved values lie disproportionately in the upper tail, so the complete-case mean is attenuated downward. This is not a precision loss and is not remedied by additional collection.

5. Benchmarking and partial correction

Table 4. Successive corrections against the external benchmark.
EstimatorValueDeviation from benchmarkCorrects for
Complete-case mean$64,153-10.6%nothing
Reweighted, board seniority mix$68,465-4.6%the disclosure filter
Reweighted, regional seniority mix$69,754-2.9%disclosure and coverage composition
Official benchmark$71,800

The standard error of the complete-case mean is $704, implying a nominal 95% interval of [$62,773, $65,534]. The benchmark lies outside it. The interval is not mis-computed; it quantifies a component of error that is close to zero by construction while the operative errors lie entirely outside its scope. Additional collection would narrow it around the same displaced center.

Reweighting to seniority recovers 56% of the deviation, and reweighting to the regional rather than the board composition a further 17 points, for 73% in total. The residual arises within seniority bands, where disclosing employers pay less than non-disclosing ones. No weighting scheme constructed from the observed data can address within-band selection, since the required distributional information is precisely what is unobserved.

Bar chart comparing $64,153, $69,754 and $71,800.
Figure 1. The scraped mean, the partially corrected mean, and the external benchmark.

6. Coverage

Table 5. Frame coverage by group, from the employer census. Overall coverage is 70.3%.
GroupAppears on a public boardRelative pay
Small employers79.2%around average
Medium employers74.3%around average
Large employers50.5%above average
Junior roles75.4%below average
Lead roles58.2%well above average

Coverage error and disclosure error operate in the same direction, both removing higher-paid roles. Concordant biases are more tractable analytically than offsetting ones, but the offsetting case is the more hazardous in practice: an estimate that appears correct while resting on two large errors of opposite sign provides no indication that investigation is warranted.

Chart of scrape coverage against the market it represents.
Figure 2. Coverage of the frame against the market it is meant to represent, and where the missing listings sit.

7. Variable-specific validity

Table 6. The same frame supports one estimand and not the other.
EstimandScrapedBenchmarkDeviationFilters applied
Mean advertised salary$64,153$71,800-10.6%coverage and disclosure
Share fully remote0.0%37.2%-100.0%coverage only

Remote status is recorded on substantially all listings, so no disclosure filter operates on it, and coverage is approximately neutral with respect to it. The dataset is therefore fit for one purpose and unfit for another. Validity is a property of the estimand given the selection mechanism, not a property of the dataset.

8. Discussion

Three conclusions generalize. First, exhaustive collection removes the error component for which standard machinery exists and leaves the others intact, so the reassurance offered by a narrow interval on scraped data is spurious. Second, the unit of analysis in scraped data is frequently not the unit of storage, and entity resolution belongs to the analysis rather than to preprocessing. Third, external benchmarking is not optional: every error quantified in this report was invisible from within the dataset.

Ethical and legal considerations attach at collection rather than at analysis. Whether collection is permitted is governed by the terms of service and by robots.txt; whether it is considerate is governed by rate limiting; and whether personal data may be retained is governed by data protection law independent of both. None of the three is enforced by the collection succeeding technically.

9. Conclusion

Mean advertised salary from the scraped frame is $64,153, deviating from the benchmark of $71,800 by -10.6%. Reweighting on seniority recovers 73% of the deviation; the remainder is attributable to within-band disclosure selection and to the 30% of vacancies never advertised publicly. Listing counts overstate vacancy counts by 71.0%. Remote-working prevalence is estimated accurately, illustrating that frame validity must be established per estimand.

References

  • Groves, R. M., & Lyberg, L. (2010). Total survey error: past, present, and future. Public Opinion Quarterly, 74(5), 849–879.
  • Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581–592.
  • Little, R. J. A., & Rubin, D. B. (2019). Statistical Analysis with Missing Data (3rd ed.). Wiley.
  • Meng, X.-L. (2018). Statistical paradises and paradoxes in big data. Annals of Applied Statistics, 12(2), 685–726.
  • Bradley, V. C., Kuriwaki, S., Isakov, M., et al. (2021). Unrepresentative big surveys significantly overestimated US vaccine uptake. Nature, 600, 695–700.
  • Christen, P. (2012). Data Matching: Concepts and Techniques for Record Linkage. Springer.
  • Salganik, M. J. (2018). Bit by Bit: Social Research in the Digital Age. Princeton University Press.

Reproducibility

The dataset (capstone-web-scraping-job-postings.xlsx), containing the raw capture with duplicates and parse failures intact, the collection plan and permissions record, and the external benchmark sheet, accompanies the chapter together with an executable notebook reproducing every statistic, table and figure. Analyses use NumPy, pandas and Matplotlib.

From Statistics, Data Science and AI: A Visual Handbook by John Fisher. Every statistic, table, and figure in this report is reproduced by the companion notebook.