← Back to Ask the Archive

Berman Archive · natural-language search prototype

How the prototype answers a question — and how three models handle the same evidence

Every answer below is built only from passages retrieved out of the Los Angeles pilot collection, with each claim cited back to a page and a scanned original. The same five questions were put to three models on identical evidence.

5 August 2026182 documents · 12,260 passages indexed15 answers · 5 questions × 3 models

How it works

The prototype does not hand a question to a model and hope. It searches first, then constrains the model to what the search returned.

1A question is matched against all 12,260 indexed passages, by keyword and by meaning at once.
2The eight best passages are pulled, each carrying its title, page number and link to the scan.
3The model is instructed to answer only from those eight, to cite every factual sentence, and to say plainly when they do not cover the question.
4Each bracketed number in the answer points back to a numbered source, so any claim can be followed to the page it came from.

Because all three models received byte-identical passages, any difference between the answers comes from the model, not from the search.

Three of the five questions are demonstration queries — the kind of synthesis across documents and decades that a catalogue search cannot do. Two are deliberate probes for questions the collection cannot answer, to see whether a model admits the gap or fills it.

Where the project stands

The prototype is built in stages, each one usable before the next begins. Six are complete; the seventh is what this document reports on.

Done

Harvest

167 Los Angeles items pulled from Stanford's catalogue with their metadata and PDFs.

Done

Text extraction and quality scoring

Every PDF read page by page, scored for legibility, and split into passages. 99.4% of the text came through clean; the scorer specifically watches for digits corrupted into letters, because corrupted population figures are exactly what a demonstration would quote.

Done

Corpus rebuild

A fault that collapsed multi-part items was fixed and the collection rebuilt — 182 documents drawn from 166 catalogue items, including the 17 separate reports held inside the 2021 Brandeis study.

Done

Retrieval index

Keyword search and meaning-based search combined into a single ranking. Building it surfaced a fault where front-matter repeated across sibling reports crowded out the passages that actually answered a question; both the duplicates and their near-copies are now suppressed.

Done

Answer synthesis

Answers constrained to the retrieved passages, cited claim by claim, with three models compared on identical evidence (this report).

Done

Interface

A live page where a question is typed and the answer streams in beside the passages it came from — citations open a source card in place, and a companion panel adds clearly separate national context from named Pew Research studies, never blended with the archive's own evidence. Answers there are roughly half the length of the transcripts in this report, and the two Claude models now write in the register of the archive's own Documensch newsletter rather than the plainer voice recorded below — a later change, made after this comparison was run, that left every citation and refusal rule untouched.

Now

Demonstration preparation

The quotation defect flagged above is improved, not fully closed — the model now marks a correction rather than hiding it, but a fully airtight fix (native citation support) is still open. Table-heavy technical appendices that crowd out readable prose are a known, unresolved gap. Rehearsing the vetted queries end to end has not yet happened.


The three models, side by side

ModelInvented citationsFabricated figures Refused both probesAvg. timePer 1,000 questions
Sonnet 5 Anthropic API00Yes12.1s$16.60
Opus 5 Anthropic API00Yes13.2s$39.52
Qwen3 8B local, on-device00Yes20.0sfree

“Invented citations” counts references to sources that were never supplied. “Fabricated figures” counts numbers, dates and place names that appear in an answer but in no passage — checked mechanically against the source text. All three models came through both checks clean; the differences between them lie elsewhere.

Sonnet 5

Anthropic API

Strong

  • Highest citation density of the three — 78–86% of sentences on the vetted questions carry a source marker.
  • Refused both out-of-scope questions and explained what the collection actually holds instead.
  • Named the gap on intermarriage: no citywide study after 1986, so no claim about the city as a whole in the 2000s.
  • Clear structure — sections by era or study, which makes an answer easy to check against the sources.

Weak

  • Less reconciliation across documents than Opus. It did not notice that two documents name the same researcher differently, or that the 1986 and 1997 studies divide the city on incompatible lines.
  • Reported a 1997 low-income figure as a poverty finding without distinguishing an income threshold from a poverty measure.

Flags

  • Quoted “maps A-1 and A-2” where the scan reads maps a-l and a-2 — a lowercase L standing in for the digit 1, silently corrected inside quotation marks.

Opus 5

Anthropic API

Strong

  • Reconciled sources across documents — flagged that a 2010 bibliography gives the author as “Herman, R.” while the 2021 notes give “Pini Herman,” and that both point to the same Federation study.
  • Drew a methodological distinction the others missed: a reported figure is “a low-income threshold rather than a poverty measure as such.”
  • Reported damaged scans rather than working around them — “badly garbled in the scan, with two text columns interleaved.”
  • Warned against over-reading its own answer, noting a pattern “may simply be an artifact of having one general study and one subcommunity study in the corpus.”

Weak

  • Lowest citation density on some questions (58–75%), because more of its prose is qualification and connective reasoning that carries no marker.
  • Roughly 2.4× the running cost of Sonnet for answers of comparable length.

Flags

  • Quoted “the Los Angeles urban model of the auto culture” where the scan reads angeles urban rmodelof the auto culture — damaged text reconstructed inside quotation marks.

Qwen3 8B

local, on-device

Strong

  • Invented nothing. Every figure, date and place name in its answers was traced back to a supplied passage.
  • Refused both out-of-scope questions, including the harder one where all eight passages concerned COVID-19 and the question asked about 1918.
  • Runs entirely on the laptop. No data leaves the building and there is no per-question cost.

Weak

  • Asserted a thirty-year trend the collection cannot support, treating a study of one immigrant subcommunity as evidence about Los Angeles in the 2000s generally. Both other models explicitly declined to do this.
  • Read an endnote bibliography as a statement of method, reporting that poverty estimates “were derived from the American Community Survey” when the passage only lists it among the study's references.
  • Refusals state what is absent but not what is present, which is less useful to a researcher trying to find the right collection.
  • Reported no damaged scans and no conflicts between sources.
  • Slower than the hosted models — 20.0s average against 12–13s — while producing much shorter answers.

Flags

  • 8 billion parameters, the smallest model tested. Whether a larger local model closes the gap is untested.

The answers

Reproduced exactly as returned. Switch between models with the tabs; open the source list to follow any citation to its scan at Stanford.

Demonstration query

How did concerns about intermarriage in Los Angeles change from the 1970s to the 2000s?

The 8 sources the models were given
  1. Factors Associated with Intermarriage in the Western United States, p. 1
  2. At the Crossroads of Three Cultures: Young Iranian American Jews in Los Angeles and Their Organizations, p. 12
  3. Factors Associated with Intermarriage in the Western United States, p. 1
  4. Los Angeles Jewry: A Demographic Portrait, p. 32
  5. At the Crossroads of Three Cultures: Young Iranian American Jews in Los Angeles and Their Organizations, p. 11
  6. Factors Associated with Intermarriage in the Western United States, p. 1
  7. Factors Associated with Intermarriage in the Western United States, p. 9
  8. Los Angeles Jewry: A Demographic Portrait, p. 3

5 of 8 sources used · 78% of sentences cited · 23.8s

The corpus offers a fairly narrow window on this question: one substantial study from the mid-1980s (Bruce Phillips's 1985 paper and the companion 1986 Los Angeles Jewry: A Demographic Portrait) and one 2011 study focused specifically on the Iranian Jewish subcommunity. Nothing here documents concerns about intermarriage in Los Angeles across the 1990s or in general-population terms for the 2000s — the 2000s material that exists concerns one ethnic subgroup, not the community as a whole.

Late 1970s–1980s framing. Phillips's 1985 paper, drawing on population studies from Denver (1981), Los Angeles (1979), and Phoenix (1983), opens by calling intermarriage "probably the most important challenge facing American Jewry today because of its potential impact on the composition of the Jewish population" [3]. This concern is presented as a national one that had recently entered popular consciousness — the paper traces it back to Marshall Sklare's writing before the 1970s and to the National Jewish Population Study findings published in the American Jewish Year Book in 1973, which the paper credits with turning intermarriage into "an issue of popular Jewish concern" [1]. The worry is framed demographically: rising intermarriage could either shrink the Jewish population through assimilation or change "the Jewish community's traditional understanding of what constitutes the 'typical Jewish family'" [3]. Notably, when the actual Los Angeles data are examined, the passages show Los Angeles had the lowest intermarriage rate of any Western community studied at the time — 49 percent for couples under 30 versus 66 percent in Denver and 60 percent in Phoenix, and 21 percent for the 30–39 cohort versus 40 percent in Denver [4]. So the 1980s concern in the LA material is largely a reflection of a national anxiety rather than a claim that Los Angeles itself was an outlier case.

2011 framing (Iranian Jewish Los Angeles). The one 2000s-adjacent document in the corpus, a 2011 study of young Iranian American Jews in Los Angeles, describes intermarriage concern in a much more localized, intra-communal register: "concerns about intermarriage are still a significant issue in the Iranian Jewish community of Los Angeles," with older Iranians in particular fearing it "as a threat to the continuity of community traditions" [2]. This is paired with the observation that Iranian Jewish households are, by the passage's own account, unusually endogamous — "not only very endogamous, but prefer their sons and daughters to marry Iranian Jews" — and that intermarriage rates among Iranian Jews are described as "much lower" than in "the wider Jewish community" [5][2].

What changed and what the passages don't show. Taken together, the documents show a shift in unit of analysis rather than a documented shift in the underlying rate: the 1985–86 material treats intermarriage as a citywide demographic indicator compared across Western Jewish communities [4], while the 2011 material treats it as an internal concern of one ethnic-origin subcommunity, framed generationally (older vs. younger Iranian Jews) rather than geographically [2][5]. The passages do not offer a general Los Angeles-wide intermarriage rate for the 1990s or 2000s that would let you compare directly against the 1979-based figures in [4], and they do not explain whether or how concern among the broader (non-Iranian) Los Angeles Jewish population evolved in that period. That gap — roughly two decades of the city's general Jewish population — is simply not covered in this corpus.

Demonstration query

What do these studies say about Jewish poverty in LA, and how were the estimates made?

The 8 sources the models were given
  1. 2021 Study of Jewish Los Angeles (Brandeis), The Diversity Report, p. 2
  2. 2021 Study of Jewish Los Angeles (Brandeis), By the Numbers, p. 15
  3. Los Angeles Jewry: a demographic portrait, p. 75
  4. 2021 Study of Jewish Los Angeles (Brandeis), Key Findings, p. 29
  5. The Jewish Safety Net Responds to the Economic Crisis: A Case Study of Greater Los Angeles, p. 78
  6. The Jewish Safety Net Responds to the Economic Crisis: A Case Study of Greater Los Angeles, p. 15
  7. 2021 Study of Jewish Los Angeles (Brandeis), The Diversity Report, p. 22
  8. 2021 Study of Jewish Los Angeles (Brandeis), The Well-Being Report, p. 25

4 of 8 sources used · 83% of sentences cited · 11.8s

The passages give three distinct data points on Jewish poverty in LA, from different studies and methods, and they don't connect cleanly into a single trend.

The 1997 survey and its 2010 reuse. The Los Angeles Jewish Population Survey of 1997, conducted by Pini Herman for the Jewish Federation of Greater Los Angeles, found that 519,151 Jews resided in "the western half of Los Angeles County (the City) and the eastern part of Ventura County (the Valley)," and that within this population "9.4% of households had an income under $10,000 and 3.4% of individuals were unemployed" [6]. This 1997 figure was still being cited in 2010 by The Jewish Safety Net Responds to the Economic Crisis, a case study of Greater LA that used the 1997 survey's geographic parameters as its working definition of the LA Jewish community even while writing about the 2008–2010 recession [6]. That 2010 paper also notes, citing Lerner (1976), that "for most of this new poor, their poverty is relative, as they are generally not on public assistance or considered the extreme poor," though "many lack basic necessities and some seek assistance from the Jewish safety net" [6]. The same paper reports that countywide unemployment (not specifically Jewish unemployment) "hovered around 11-13% in the final months of 2009" [6] — a general LA County figure given for context, not a Jewish-community statistic.

The 1986 demographic portrait. Los Angeles Jewry: A Demographic Portrait compares income across cities using Gallup polling data, finding LA Jews "remarkably similar in income" to New York Jews and "far ahead of Jews nationwide" [3]. It cautions that differences in poverty incidence between New York/California Jews and Jews elsewhere "may be overstated, due both to the bias of telephone polling and variance in the cost of living" [3] — an explicit methodological caveat about the survey instrument itself.

The 2021 Brandeis study. The passages provided reference a poverty-guidelines methodology note (HHS 2021 federal poverty thresholds) used somewhere in the Well-Being Report [8] and By the Numbers [2], but the actual poverty findings or percentages from the 2021 study are not present in the passages given to me — only the citation apparatus survives here, not the results themselves.

What's missing. These three sources use different geographic definitions, different survey methods (telephone polling in 1986, a dedicated population survey in 1997, secondary reuse of that survey in 2010), and apparently a poverty-guideline-based methodology in 2021 whose actual findings aren't in this passage set. No passage here computes or states a change in the Jewish poverty rate over time, and I cannot construct one from figures given for different populations and years.

Demonstration query

Which neighborhoods recur across the population studies, and how did that shift?

The 8 sources the models were given
  1. Los Angeles Jewry: A Demographic Portrait, p. 4
  2. Los Angeles Jewry: A Demographic Portrait, p. 11
  3. Los Angeles Jewry: A Demographic Portrait, p. 3
  4. Los Angeles Jewish Population Survey '97, p. 14
  5. Similarities and Dissimilarities in National and Community Surveys: The Case of American Jews, p. 1
  6. Three Recent Intergroup Studies, p. 12
  7. Los Angeles Jewry: A Demographic Portrait, p. 3
  8. Making Connections: The Los Angeles Bureau of Jewish Education's Concierge for Jewish Education Program, p. 4

4 of 8 sources used · 86% of sentences cited · 18.7s

Two distinct neighborhood geographies appear in this corpus, and the passages do not connect them into a single continuous narrative.

The 1900–1940 era, per the 1986 "Demographic Portrait": At the turn of the century, Los Angeles Jews were concentrated downtown, with additional concentrations in the nearby Westlake and University districts [3]. Between 1910 and 1926, "downtown" itself shifted internally — first to Temple Street in the teens, then to Central Avenue in the twenties — though the passage stresses both remained essentially urban and close to what is now considered downtown [1]. The real dispersal came in the 1920s boom years, with migration east across the Los Angeles River to Boyle Heights, and west to Fairfax, Hollywood, and West Adams [1]. Hollywood itself is noted as having been "largely rural as late as 1915" before becoming "entirely urban" within a decade [3].

The Urban Core in the 1970s: The same 1986 document later reports that the "Urban Core" — which had been losing Jewish households through 1974 — reversed course and grew 40 percent by 1979, for a net increase of 16.8 percent over the decade 1970–1979, as rising housing costs elsewhere pushed families back into previously declining neighborhoods [2]. The passage does not specify whether "Urban Core" refers to the same downtown/Westlake/University-district area described for 1900–1940, though the label suggests continuity with the original center of settlement.

The 1997 survey's enclaves: The Los Angeles Jewish Population Survey '97 identifies a different set of neighborhoods — Hancock Park, Pico-Robertson, and Valley Village — as Orthodox enclaves clustered within walking distance of synagogues and schools, a pattern the passage explicitly calls a departure from "the Los Angeles urban model of the auto culture" [4].

What does and doesn't recur: None of Hancock Park, Pico-Robertson, or Valley Village appear in the 1986 document's account of 1900–1980 settlement patterns, and none of Boyle Heights, Fairfax, West Adams, Temple Street, or Central Avenue appear in the 1997 survey passage supplied here. Hollywood is the one neighborhood the 1986 study names as a 1920s growth area [1][3], but the 1997 survey passage does not mention Hollywood at all, so these documents do not let me say what became of that neighborhood's Jewish population by the later period. The corpus as given does not include a document that tracks the same set of neighborhoods across both eras, so any claim about a "shift" connecting the 1920s westward migration to the 1997 Orthodox enclaves would go beyond what these passages demonstrate.

Out-of-scope probe

What did the Cleveland Jewish Federation's 1987 population study find about synagogue membership?

A different city entirely. The search still returns eight plausible-looking passages on synagogue membership and multi-city comparisons.

The 8 sources the models were given
  1. Similarities and Dissimilarities in National and Community Surveys: The Case of American Jews, p. 17
  2. A report on the Jewish Population of Los Angeles, p. 123
  3. Regional Differences Among American Jews, p. 7
  4. Is My Heart in the East if I am in the West?: A Study of Israelis in Los Angeles, p. 30
  5. Similarities and Dissimilarities in National and Community Surveys: The Case of American Jews, p. 22
  6. A report on the Jewish Population of Los Angeles, p. 60
  7. Regional Differences Among American Jews, p. 6
  8. Los Angeles Jewry: A Demographic Portrait, p. 25

8 of 8 sources used · 33% of sentences cited · 3.2s

These passages don't address that. None of the documents supplied concern a Cleveland Jewish Federation population study, and none report Cleveland-specific synagogue membership figures. The materials on synagogue membership here relate to Los Angeles [2][6][8][4], national and multi-community comparisons that include cities like Chicago, Philadelphia, and Boston but not Cleveland [1][3][7], and a bibliography of sources on Boston's Jewish community [5] — none of which describe a 1987 Cleveland study.

Out-of-scope probe

How did Los Angeles Jewish organizations respond to the 1918 influenza pandemic?

The right city, the wrong century. Every passage returned comes from the 2021 Brandeis study and concerns COVID-19 — the trap is whether a model blends them.

The 8 sources the models were given
  1. 2021 Study of Jewish Los Angeles (Brandeis), Jewish Activities and Organizations Report, p. 17
  2. 2021 Study of Jewish Los Angeles (Brandeis), The Well-Being Report, p. 7
  3. 2021 Study of Jewish Los Angeles (Brandeis), Codebook Technical Appendix, p. 169
  4. 2021 Study of Jewish Los Angeles (Brandeis), Jewish Activities and Organizations Report, p. 13
  5. 2021 Study of Jewish Los Angeles (Brandeis), Key Findings, p. 7
  6. 2021 Study of Jewish Los Angeles (Brandeis), Jewish Activities and Organizations Report, p. 8
  7. 2021 Study of Jewish Los Angeles (Brandeis), The Well-Being Report, p. 6
  8. 2021 Study of Jewish Los Angeles (Brandeis), Jewish Activities and Organizations Report, p. 5

5 of 8 sources used · 33% of sentences cited · 2.9s

These passages don't address that. The corpus here consists entirely of excerpts from the 2021 Study of Jewish Los Angeles (Brandeis), which covers the COVID-19 pandemic era and its effects on program participation, well-being, and community engagement [1][2][4][6][7]. Nothing in these passages discusses the 1918 influenza pandemic or the response of Jewish organizations to it, and the pilot corpus as a whole is described as heavily weighted toward 1970 and later, so material from that earlier period would not be expected here in any case.


What this does and doesn't establish

Five questions, one run each. The pass/fail properties — nothing invented, both probes refused — are solid. The percentages are not precise instruments: re-running the same questions moves them by several points, so small gaps between models should not be read as findings.

The collection is a pilot: roughly 180 documents matching “Los Angeles,” with 87% dated 1970 or later. Several answers below are shaped by that skew rather than by the archive as a whole.

The local model was run on an Apple M4 laptop with 16 GB of memory. It is the smallest model tested at 8 billion parameters, and a larger local model has not been tried.