J-SOC public research package

CONTENTS
Full report and five-page brief; financial records in CSV/JSON; source register
in CSV/JSON; selected unchanged original documents; SHA256SUMS.txt.

FORECAST COLLECTION
2,336 forecast listings, each with its own ID, from 70 named report files.
Collection credited to Wallface: https://x.com/Wallface
CSV/JSON exports retain printed fields, report filename, page and workbook row.
This is not a count of unique events, or of events confirmed to have happened.
Listings dated in the future are forecasts of planned events, not events that happened.
Leaving out the 77 earlier-edition listings gives 2,259, still not deduplicated.
Risk labels are the source's own assessments, not ours.
original_url links each listing to its report on gov.il (all 70 located).
j-soc-forecast-reports.json gives each original's SHA-256 and page count, as
as we downloaded them, so a copy can be checked against ours.
Per-report counts match the workbook. Transcription check: 100 listings drawn
with Python random.Random(20260928).sample(records, 100) and compared with the
originals; 1 error found (F37-E013). 95% CI for the error rate: 0.03%-5.4%.
named_groups: organisers and posting or exposure accounts named in the source.
A name in this field does not by itself mean the account organised the event.
notes: source notes, plus editor's notes (marked as such) where a listing was
corrected against the original report.
Parsed numeric follower totals and derived coordinate/distance fields are omitted
from the public export because those calculations have not been validated.
Printed exposure, coordinates and proximity text are retained, not endorsed.
The source workbook SHA-256 and detailed provenance are in j-soc-forecasts.json.
CSV text beginning with a formula prefix is escaped for spreadsheet safety.
Private individuals' names are withheld ("individual (name withheld)");
organisations and public figures are kept. description_hebrew is omitted because
Hebrew-script names cannot be reliably matched. The linked originals are unedited.

FINANCIAL DATA DICTIONARY
id: stable research record ID, not an official identifier.
payer / recipient: labels for the documented relationship.
amount / currency: numeric value and ISO currency, without conversion.
measure: order value, ceiling, schedule, invoice, receipt or disbursement.
period: source reporting label or document date, not inferred cash date.
order_id / record_locator: official order or filing locator.
source_ids / source_urls: join to the source register and original links.
note: interpretation and material limits.
review_status: how the record was checked for this edition.

SOURCE REGISTER
Original URLs, publication names, locators, source type, review/access notes,
preserved-copy paths and SHA-256 hashes. A preserved_copy value is relative to
the package root. Empty means no original is bundled, not that no copy exists.
Source IDs are stable within this edition and match the page and reports.

METHOD
Do not sum financial rows: budgets, cumulative snapshots, invoices and cash
declarations can overlap. Do not merge ILS with USD. Company and order IDs are
text identifiers. Workbook timestamps do not determine accounting cutoffs.
Amounts from Ministry workbooks come from earlier copies and are labelled so.
Workbook figures that could not be traced to a recoverable source are omitted.
Account selection, citation and relay are not proof of financial relationships.
Original evidence is reproduced unchanged. Secondary-hosted copies are labelled.

INTEGRITY
SHA256SUMS.txt lists package files other than itself. The website version also
lists the ZIP hash. Compare hashes after download to check for changed bytes.
