fix: run usp_Database_Load and fix bulk insert scaling in load_sql.py - #294
Closed
benhayes21 wants to merge 2 commits into
Closed
benhayes21 wants to merge 2 commits into
benhayes21 wants to merge 2 commits into
Conversation
The EXEC usp_Database_Load call was commented out, so staging tables loaded but the full database load never ran. Uncommented it and switched to engine.begin() so the call actually commits instead of being rolled back on connection close. Also replaced method='multi' with fast_executemany=True, which batches inserts at the pyodbc driver level instead of building one big multi-row INSERT bounded by SQL Server's 2100-param limit, so it scales to OED input files with millions of rows. Account dates are normalized to ISO 8601 per the OED spec before loading. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Build PreviewYou can find files attached to the below linked Workflow Run URL (Logs).
|
load_sql.py was truncating location.csv to 10 rows via a leftover .head(10) debug call, so only a handful of locations ever reached staging regardless of input size. Separately, LocationDetail was defined in the schema but never populated: usp_Database_Load only ran usp_Location_Load, which inserts the 4-column Location skeleton, with no procedure loading the ~150 detail attributes (address, geocoding, occupancy/construction codes, vulnerability fields, etc.) that land in _import_location. Added usp_LocationDetail_Load, which joins _import_location to _businesskeys_location on the natural key to resolve LocationId and inserts the detail columns, and wired it into usp_Database_Load. Verified against a local SQL Server instance: full 12598-row location.csv now loads end-to-end into Location and LocationDetail with no orphaned or duplicate rows. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Build PreviewYou can find files attached to the below linked Workflow Run URL (Logs).
|
benhayes21
added a commit
that referenced
this pull request
Sep 17, 2026
Folds in the fixes from #294 on top of this branch's relational- schema rework: - load_sql.py: batch to_sql() writes with chunksize=10000 (works with fast_executemany to scale past SQL Server's 2100-param limit on large OED input files) and normalize PolInceptionDate/PolExpiryDate to ISO 8601 before staging, since the OED spec requires it and source files may use a locale-specific date format. - LocationDetail was defined in the schema (and even cleared by usp_Database_Load's idempotency reset) but never populated: no procedure existed to load its ~160 detail attributes from _import_location. Added usp_LocationDetail_Load, joining _businesskeys_location to _import_location on the natural key (matching usp_Location_Load's existing join pattern) and wired it into usp_Database_Load right after usp_Location_Load. LocGroup is intentionally excluded from the column list since this branch already extracts it into the separate LocationGroup table. Verified against a local SQL Server instance: full 12598-row location.csv loads end-to-end into Location and LocationDetail with no orphaned or duplicate rows, and correct attribute values. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #293
Summary
Test plan