Repository navigation
Conversation
…edicated module account_invoice_import_invoice2data Update README.rst and headers to latest OCA conventions.
Better key names in the parsed_inv dict parsed_inv doesn't need to be JSON serializable anymore (small drawback: the invoice is parsed a second time on the second step... but the second step is rarely used)
Move code from account_invoice_import_invoice2data to account_invoice_import
Update REAME and some interface strings about UBL being an ISO standard Small code changes
…voice dict, cleaner organisation) Code refactoring: move code in base_business_document_import, factorise code for tax matching (it was duplicated in UBL and ZUGFeRD) Now support PDF with embedded UBL XML file Enable unittests on account_invoice_import_ubl More absolute xpath in account_invoice_import_ubl instead of relative xpath WARNING: these are big changes, I may have broken a few details
…te dir and the built-in templates Also allow to use only a local template dir README updated to explain how to configure all this
Special thanks to Sébastien Beau for his help to achieve this
Add support for partner bank matching on invoice update (before, it was only supported on invoice creation)
[FIX] LINT Use try/except when importing external libs Remove self.ensure_one() that has nothing to do in an api.model method
…a recent version of pdftotext on travis's ubuntu 12.04 images
Rename __openerp__.py to __manifest__.py and set installable to False
Also port all the modules that generate the XML documents: account_invoice_ubl, account_invoice_zugferd, purchase_order_ubl and sale_order_ubl
… module Fix spelling mistake and other remarks on README by Tarteo
* Update to work with latest version of invoice2data * Add requirements.txt file
Update the account_invoice_download_weboob following the changes in account_invoice_download Add README.rst for the 2 modules
Update text displayed in the invoice import wizard and make list of supported formats modular (like in bank statement import)
Updated by Actualizar ficheiros PO com o novo POT (msgmerge) hook in Weblate.
bosd
force-pushed
the
16.0-store_invoice2data_templates
branch
from
July 18, 2026 17:36
007b588 to
446ff7c
Compare
DB-stored invoice2data templates + GUI builder as a new addon that depends on account_invoice_import_invoice2data (this PR is stacked on top of OCA#1220). - `invoice2data.template` model with two authoring modes: guided (name + keywords + `field_ids` o2m composing the JSON at read time) and power- user (paste a full invoice2data JSON template into a dedicated tab). - Per-field `replace` pair (issue invoice2data#497) and `extract_number` flag (issue invoice2data#652) exposed as first-class columns. - Form actions: Preview text, Test (full extract_data() vs chatter PDF), and Suggest fields (wires the lib's suggested_template + label detection to pre-fill the field grid). - Wizard extension reimplements invoice2data_parse_invoice: builds the templates list via a `_invoice2data_collect_templates` hook that concatenates disk + DB. Uses `with NamedTemporaryFile(...)` + flush() (matches the review on OCA#1220). - Security: dedicated 'Manage invoice2data templates' group. - Tests: JSON round-trip, structured composition, replace pair, extract_number opt-in semantics, type/active filtering, wizard merge. Selection lists are driven from invoice2data.extract.schema so adding a new canonical field upstream automatically becomes selectable here.
added 2 commits
July 19, 2026 21:48
Two red-CI fixes: - `_field_selection` was defined as a zero-arg function but Odoo passes `self` to selection callables; test errors were: `TypeError: _field_selection() takes 0 positional arguments but 1 was given` - black + isort reformats to satisfy OCA/edi's pre-commit config.
Pre-commit's 'unchecked-in-files' guard was failing because the OCA
setup/<addon>/{setup.py, odoo/addons/<addon>} scaffolding wasn't
committed alongside the module. Add it (matches what
setuptools-odoo-make-default generates).
bosd
force-pushed
the
16.0-store_invoice2data_templates
branch
from
September 1, 2026 16:21
8cb94d4 to
fb69c75
Compare
Contributor
Author
|
force pushed to recreate runboat ⛵ |
added 6 commits
September 1, 2026 18:47
…g-template warning - security: auto-grant the template-manager group to base.group_system so a fresh admin sees the Create button. Drop the group's accounting category so it lands under 'Other' in the user form (developer-mode visible) instead of as a dropdown option in the Accounting section. - models.action_suggest_fields: drop the erroneous 'name=' kwarg from suggested_template(); the invoice2data API signature is suggested_template(text: str), not (text, name=...). Was raising TypeError on button click. - models.action_test: when result['template_name'] does not equal self.name, emit a warning so the author sees that the extracted fields came from a different (shipped) template rather than the one they were authoring -- otherwise a plausible-looking test result masks the fact that their own template did not match at all.
…of Test results - Split the mixed 'Test results' notebook page: 'Preview text' now has its own tab (populated by the Preview text button), 'Test results' keeps warnings + JSON only. - Warning banner moves out of the <group> wrapper so it renders full-width at the top of the page rather than in one column of a two-column layout.
…e disk templates Two motivations came out of downstream testing on Runboat: 1) The Test button was matching against the disk template pool, which meant a shipped template (e.g. com.AzureInterior.yml) intercepted a PDF the author was iterating on — the results shown had nothing to do with the template being authored. (The 'matched != self.name' warning added in the previous commit surfaces this, but doesn't help an author who wants the disk pool out of the way while iterating.) 2) An author working on a live install may want disk templates OFF globally (they have DB templates covering the same ground) or ON (they want disk templates to fill the gaps). Changes: - New ir.config_parameter 'account_invoice_import_invoice2data_db_templates.include_disk_templates' (default 'True'). action_test consults it: when False, extract_data runs against active DB templates only, no disk pool. - New action_test_isolated method exposed as a second header button, gated on base.group_no_one so it only appears in developer mode. Runs against ONLY [self], regardless of the config parameter — the most focused feedback loop for iterating on one template. - action_test refactored to call a shared _run_test(include_disk, isolated); no behavioural change on the default path when the setting is True.
…tags The Text field (one keyword per line) was serviceable but did not autocomplete, made cross-template reuse invisible, and required scrolling even for the common 1-3 keyword case. Move both fields to Many2many pointing at a shared invoice2data.template.keyword pool with the many2many_tags widget. Model changes: - New invoice2data.template.keyword: minimal Char record (name, unique). - keywords / exclude_keywords on invoice2data.template become Many2many with distinct relation tables so the same keyword can be included by one template and excluded by another. - _compose_template_dict() renders the m2m .name attributes into the same list<str> shape the invoice2data lib expects at the top level of the composed template dict. View: widget='many2many_tags' with quick-create enabled so authors can type-and-Enter to add a new tag. Security: read-only for account.group_account_invoice, full CRUD for group_invoice2data_template_manager on the new keyword model. Migration to 16.0.1.1.0: - pre-migration renames the old Text columns (keywords_text_deprecated / exclude_keywords_text_deprecated). - post-migration reads them, splits on newlines, de-dupes into the keyword pool, populates both relation tables, then drops the deprecated columns. Manifest version bumped to 16.0.1.1.0.
Gives an author a starting point: read the bundled invoice2data templates (read_templates()) into invoice2data.template DB records they can then edit / delete / customise. Wizard shape (invoice2data.template.import.disk.wizard): - name_regex: Python regex on the template's filename; blank = all (currently ~215 bundled templates). Example: '^nl\.' for the Dutch subset, 'shell' for anything shell-branded. - mode: skip_existing (default) | overwrite. Existing-by-name are either left alone or full-overwritten. - dry_run: reports counts without touching any records. - result_summary populated after import so the user sees considered / created / updated / skipped counts and any errors. Import strategy: serialise the disk template dict verbatim into the record's 'template' JSON blob so nothing (options, lines blocks, options.date_formats, etc.) is lost through a lossy round-trip; the authoritative-JSON path already exists on the template model. Keywords and exclude_keywords are ALSO populated in the m2m tags so they show up in the record header. Menu action: Accounting > Configuration > Import disk templates (sequence 81, sibling of the templates menu at 80). Security: manager group only (create+read+write+unlink on the transient model).
Pre-commit on the previous push (c2c2199) reported one prettier reflow: security/invoice2data_template_groups.xml had a multi-line <field name=... eval=... /> that prettier wanted on a single line -- reflow it. Also runs black over the four new/modified .py files in the current stack so the next push is pre-commit-clean.
Merged
2 of 3 tasks
added 7 commits
September 1, 2026 21:19
The many2many_tags widget on keywords / exclude_keywords was rendered as Many2ManyTagsFieldColorEditable (the color-editing variant) because 'color_field' was present in the options -- Odoo 16 (OWL) rejects the prop with 'colorField is not a string' the moment a form using the field opens (e.g. clicking New Template): OwlError: Invalid props for component 'Many2ManyTagsFieldColorEditable': 'colorField' is not a string The keyword model has no color field and doesn't want color editing, so drop the option entirely. Odoo then binds the plain Many2ManyTagsField component, which validates and renders cleanly. Also add the new OCA logo (aa7c7fc..., orange sun-burst over 'OCA' on dark blue) as static/description/icon.png so the module card in Apps matches sibling OCA modules.
…lds override
Two related iterations after downstream testing:
1) The Test button lumped 'no keyword match' and 'keywords matched but
extraction incomplete' into a single 'invoice2data did not match this
PDF' warning. An author whose template only has keywords (no field
regexes yet) sees 'no match', can't tell whether their keywords are
wrong or their (missing) regexes are, and doesn't know that the
library actually did engage their template.
Capture the invoice2data logger during extract_data, parse for the
'Template: X | Keywords matched.' and 'extraction was incomplete: ...'
message shapes, and surface: templates-in-pool count, extracted-text
length, set of templates whose keywords matched (up to 10), and any
incomplete-extraction reasons the lib reported per-template.
2) invoice2data's required_fields default ('date, amount, invoice_number,
issuer') is invoice-shaped and rejects extractions from non-invoice
document types (waybills, delivery notes, receipts). The lib already
supports a per-template 'required_fields:' key that overrides the
default; expose it on the DB template as a Char whose comma-separated
tokens flow into _compose_template_dict. A bare comma disables the
check entirely (empty list).
Manifest version bumped to 16.0.1.2.0 to force an Apps upgrade so the
new required_fields column is added to invoice2data_template.
…iagnostics Two related fixes: 1) The previous _run_test relied on capturing invoice2data's own log records to diagnose 'keyword-matched but incomplete extraction'. That silently loses records when a child logger (invoice2data.api, etc.) has its own level set above DEBUG -- which is Odoo's default for third-party libraries -- so the author saw only the header line and no per-template reason. Rewrite as direct per-template inspection: walk the pool in priority order, call prepare_input + matches_input, then attempt extract() and record the outcome. Emits one warning line per matched template with the actual exception message (RequiredFieldsMissingError etc.). 2) When the extraction ran through the exception-catching path, last_test_result / last_test_warnings from a previous run stayed visible. Clear both up front so the form always shows fresh state even when something silently short-circuits. Also tightens the header line to expose the isolated / include-disk mode plus text length + pool size, and adds a tip pointing to the 'Required fields = ,' escape hatch when nothing extracted.
…n invoice-shaped form Item 1 of the MVP→v1 plan (the internal design brief section 1). Before: Test / Test isolated dumped a JSON blob into last_test_result. Non-technical template authors could not read it. After: those buttons still populate last_test_result / last_test_warnings (diagnostics tab is unchanged), and now also open a modal wizard that renders the extraction as an invoice — partner block on the left, invoice header on the right, invoice lines in a page, tax lines in a sibling page, totals in the footer. Adds three transient models (invoice2data.template.preview, .preview.line, .preview.tax) and one form view. The main preview model holds mirrored account.move-shaped fields: - Partner block from partner_name/vat/email/street/zip/city/country_code with a best-effort res.partner lookup (VAT first, name fallback, silent on error). Non-matching partner is shown blank so the author sees the import would create/reject. - Header from invoice_number/date/date_due/payment_reference/currency (currency looked up by ISO code). - Amount footer from amount/amount_untaxed/amount_tax (Monetary). - Lines from result['lines'] via _line_vals (canonical name/product/ qty/uom/price_unit/price_subtotal). - Tax lines from result['tax_lines'] via _tax_line_vals. - Diagnostics tab with the missing-canonical-fields hint and the full _run_test warnings list. - Raw JSON tab (ace/json widget) for power users — the JSON blob was removed from the notebook page it used to live on so that page only shows the diagnostic banner now. Manifest bumped to 16.0.1.4.0 (new tables → module upgrade required). ir.model.access.csv grants group_account_invoice CRUD on the three transient models.
…(AI) button Three interlocking additions from the MVP→v1 plan: 1. input_module Selection (pdfium / pdftotext / pdfminer / pdfplumber / text / tesseract / ocrmypdf / gvision / docTR / paddleocr) — per- template backend pin, already supported by invoice2data core; the template author can force a specific backend when a vendor's PDF fails to parse under the site-wide cascade default. Emitted into the composed template dict as 'input_module' when set. Will also be written automatically by the future click-to-suggest wizard (the design brief's cross-backend-drift guard). 2. New 'Refresh candidates' header button + Candidates tab (candidates_summary Text). Runs find_candidates() and find_labeled_fields() on the attached PDF's extracted text and lays them out in a monospaced diagnostic table. Surfaces the WHY behind Suggest fields' proposals (e.g. 'VAT was recognised via a Dutch BTW label at offset 234-289'). 3. New 'Suggest fields (AI)' header button. Calls invoice2data's AI-1 generate_template() when the [ai] extra is installed and a provider is configured (INVOICE2DATA_AI_PROVIDER / _BASE_URL / _MODEL env vars). Missing extra -> UserError with install instructions. Adds only rows for fields not already in the grid; emits a warnings-tab banner reporting how many were added and reminding the author to Test the result. Manifest bumped to 16.0.1.5.0 for the schema addition (input_module, candidates_summary columns).
Item 2 of the MVP→v1 plan. Port of the CLI's _interactive_template flow (invoice2data/__main__.py) into an Odoo modal, adapted for the non-technical persona: instead of prompting one field at a time, show every proposed field in a single editable tree so the user reviews the whole draft at once, ticks Keep/Edit/Skip per row, and Apply commits only the kept rows to the template's field_ids. New transient models: - invoice2data.template.field.walk.wizard (template_id, source, line_ids) - invoice2data.template.field.walk.line (field_name, captured_value, proposed_regex editable, already_in_template flag, decision Selection). Wizard picks its draft from either the deterministic authoring API (suggested_template) or the AI-1 generator (generate_template) via a 'source' Selection. Rows the template already has are pre-set to Skip so users don't accidentally overwrite manual work. Empty regexes on Kept rows are demoted to Skip with a log entry (a template with a kept-but-empty-regex row would be unusable). The wizard writes to last_test_warnings after applying so the author sees a X added / Y replaced / Z skipped summary on the template form. Two new header buttons on the template form: - 'Guided suggest' (deterministic; the recommended entry point) - 'Guided suggest (AI)' (uses AI-1; requires [ai] extra + provider env vars — same UserError-with-install-instructions as the existing Suggest fields (AI) button) The plain 'Suggest fields' (bulk one-shot) button is kept as an advanced/power-user path; help text updated to make the distinction clear. Manifest bumped to 16.0.1.6.0.
…harness
Item 6 (first slice) of the MVP→v1 plan, implementing the revised
design brief (the internal design brief). Ships
the server side end-to-end today; the OWL PDF viewer that replaces
coord entry lands in a follow-up PR (design is stable).
Server pipeline (models/pdf_click.py):
- pypdfium2-based; no poppler binary required.
- suggest_at(pdf_path, page_idx, x_pt, y_pt) resolves a click:
textpage.get_index → bbox line via get_charbox walk →
find_labeled_fields / find_candidates on that one line →
field_regex_from_candidate → preview_field validation.
- suggest_area(pdf_path, page_idx, rect_pt) resolves a drag:
same pass on get_text_bounded(rect); emits {area + regex} when the
crop contains a label / typed candidate, or a position-only
{area: (?s)(.+)} clearly labelled 'breaks if the vendor moves this
block' when it doesn't.
- pdftotext cross-check when the binary is present: returns 'same' /
'no-match' / 'unavailable' for the wizard's drift badge (design brief §Q2 (b)).
- render_page_png helper for the future OWL viewer (via
page.render(scale=dpi/72).to_pil()).
Canonical field allow-list (invoice2data.field.key, data XML):
- 18 seed rows across header / amount / partner / identifier kinds.
- l10n glue modules extend by declaring more <record> entries — no
library change needed for SIREN, OIN, KvK, etc. (design brief §Q3).
Field-row hint bookkeeping (invoice2data.template.field additions):
- hint_page, hint_x, hint_y, hint_w, hint_h — reopen the click position.
- sample_line, sample_value — Test regression check
(design brief §Q4: Odoo-only, does NOT export to YAML).
UI harness (wizard/pdf_click_suggest.{py,xml}):
- 'Click to suggest' header button (developer-mode-only for now, since
it's a coord-entry harness rather than a click-on-image UX).
- Enter page + (x, y) or drag rect → Preview shows proposal, captured
value, bbox line, pdftotext status badge, and the canonical-field
dropdown (from invoice2data.field.key so the user can override).
- Apply appends the proposal as a new field row on the template,
writes the hint columns, and pins input_module='pdfium' on the
template (cross-backend-drift guard — never overwrites an explicit
non-pdfium choice).
Manifest bumped to 16.0.1.7.0 for the schema additions (hint columns,
new models).
bosd
added a commit
to invoice-x/invoice2data
that referenced
this pull request
Oct 4, 2026
…761) The value patterns in _VALUE_PATTERNS get embedded into every regex the CLI builder / db_templates Suggest Fields wizard generates for a matched identifier candidate. They were loose: 'iban': r'[A-Z0-9 ]+', 'vat': r'[A-Z0-9]+', 'bic': r'[A-Z0-9]+', A suggestion built on a valid candidate is therefore fine on the sample document (the candidate detector's own regex is strict, and the label anchor picks the right value there), but when the same suggested template is later applied to a fresh PDF the loose pattern grabs any adjacent uppercase-alphanumeric blob near the label -- an order code 'ORDER12345' can match 'bic' shaped just fine, so downstream consumers (Odoo's base_business_document_import) reject the field and the import fails. Tighten each pattern to mirror the strict validators in validators.py: - BIC: ISO 9362 shape [A-Z]{6}[A-Z0-9]{2}(?:[A-Z0-9]{3})? (8 or 11 chars, first 6 letters). - VAT: 2 alpha country + 8-14 alphanumerics. - IBAN: country + 2 digits + 11-30 alphanumeric body, tolerating embedded single-spaces/hyphens. Regression test in tests/test_template_builder.py covers the BIC case end-to-end (suggested template on a doc containing both an order code and a valid BIC -> regex captures ONLY the BIC). Reported downstream while iterating on the account_invoice_import_ invoice2data_db_templates Suggest Fields wizard (OCA/edi#1365). Co-authored-by: bosd <5e2fd43-d292-4c90-9d1f-74ff3436329a@anonaddy.me>
… / isort / prettier) + real-bug fixes CI pre-commit on the previous push flagged two real bugs plus a cascade of auto-formatter drift. Both bugs addressed and the formatters applied with OCA's pinned versions (black 22.8.0, pyupgrade 2.38.2): - F821 undefined name '_full_text' in pdf_click.py: an aborted-refactor 'if False' branch left a reference to a non-existent helper. Replace with a direct call to _pdfium_full_text(pdf_path). - W8113 attribute-string-redundant on invoice2data.template.preview.line.name: pylint-odoo flags the string= argument as redundant. The user-facing 'Description' label is still the right label on a line preview; suppress with a localised disable=attribute-string-redundant. - autoflake dropped unused imports (odoo, SUPERUSER_ID import order). - pyupgrade re-flowed the pdf_click.py docstring to a raw string. - black 22.8.0 reflowed the two migration cr.execute() blocks and multi-line import tuples in the wizard. - prettier (with plugin-xml) reflowed four view/wizard XML files to its preferred shape.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Depends on #1220 (the 16.0 migration of
account_invoice_import_invoice2data). Until that lands the diff includes its commits; once it merges this rebases cleanly ontoOCA/edi 16.0.What this adds
A new addon,
account_invoice_import_invoice2data_db_templates, that stores invoice2data templates in the Odoo database and merges them into the import wizard alongside the existing disk-loaded ones. Implements the GUI template builder the parent module's 2017 TODO list asked for.Authoring modes
field_idso2m; JSON is composed at save time.The canonical field-name selection list is driven from
invoice2data.extract.schema— adding a new canonical field upstream automatically becomes selectable here, no parallel list to maintain.Buttons on the form
to_texton the latest chatter attachment.extract_data()against the attached PDF using this template + the disk-loaded ones; surfaces the parsed dict and required-fields warnings.invoice2data.extract.template_builder.suggested_template+ the label-detection helpers to pre-fill the field grid from the latest attached PDF (the "guessing framework" the 2017 TODO list asked for).Wizard merge
Extends
account.invoice.importto reimplementinvoice2data_parse_invoicecleanly:with NamedTemporaryFile(...) as fileobj:+fileobj.flush()(matches my review on [MIG] account_invoice_import_invoice2data: Migration to 16.0 #1220)._invoice2data_collect_templateshook (disk + DB; downstream modules can override).First-class support for the 1.0 lib options
Per-field
replace(issue invoice2data#497) andextract_numberflag (issue invoice2data#652) are exposed as columns on the o2m, not buried in JSON.Security
Manage invoice2data templatesgroup on top ofaccount.group_account_invoice.ir.model.access.csvfor both the parent model and the o2m.Tests
Cover JSON round-trip, structured composition, the
replacepair,extract_numberopt-in semantics, type/active filtering, and the wizard's merged template list.External dependency
invoice2data >= 1.0(now on PyPI — published today). The 1.0 cut includesordered_load,suggested_template, the per-fieldreplace, and theextract_numberfield option this module surfaces.Roadmap (in
readme/ROADMAP.rst)area:-style templates (ties into thecamelot/Excalibur path).--new-template --ai) wired as an action.Reviewers
cc @alexis-via @MarwanBHL — happy to rebase / split / squash per your preferences.