[general]
name=Stratified Packager
description=Exports vector layers to multiple geopackages split by attributes or spatial intersection
category=
hasProcessingProvider=yes
about=Partitions the open project’s layers against a stratification layer — by attribute (following project relations) or spatially — and publishes one zipped GeoPackage per stratum, optionally with an embedded per-stratum QGIS project, layer styles, metadata and referenced style assets. Runs from the Processing Toolbox and from qgis_process.
icon=resources/images/icon.svg
tags=export,geopackage,gpkg,stratification,stratify,partition,split,divide,project
server=False

# credits and contact
author=Ivan Donisete Lonel
email=ivanlonel91@gmail.com
homepage=https://ivanlonel.github.io/stratified-packager/
repository=https://github.com/ivanlonel/stratified-packager
tracker=https://github.com/ivanlonel/stratified-packager/issues/

# QGIS context
deprecated=False
experimental=True
# plugin_dependencies=
qgisMinimumVersion=3.40
qgisMaximumVersion=4.99

# versioning
version=0.5.0
changelog=
 Version 0.5.0:
 - Layer variables: **a layer can now ride in the packaged project without being packaged itself**, through a new `matching_method` value, `project_only`. It is for a layer defined *over the delivered package* rather than over any source data — typically an OGR layer whose `|subset=` is a `SELECT` joining the tables the run writes, styled and configured in the source project. Such a layer points at a GeoPackage that does not exist until the package is built, so it is broken on the packaging machine, and that aborted the whole run: every packaged vector layer is cloned during analysis, and cloning rebuilds a layer from its URI. A `project_only` layer is now never read, never staged and gets no table of its own; in each stratum's embedded project, everything before the first `|` of its data source is replaced with that stratum's GeoPackage and every uri option after it is kept verbatim, so it lands on the recipient's machine pointing at a relative `./<stratum>.gpkg`. The columns its query compares with `=` are indexed just as a live virtual layer's are, and it is written without a computed extent for the same reason — both of those are what keep such a layer from costing minutes per draw. Two guards fail the run up front rather than shipping a package quietly missing the layer: the provider must be `ogr` (replacing a database connection string wholesale would leave a valid layer over the wrong table), and every table the query reads must be one the run creates, which is what a renamed layer would otherwise break in silence.
 - Reporting: **layers riding only in the embedded project now appear in both reports.** Remote basemaps, annotations and live virtual layers ship in every package, yet had no row in the run report or in any `report.csv` — the per-zip column contract already described their empty `path_in_zip`, and nothing ever emitted one. They now report the new `project-only` status, with an empty feature count, in the run report and the per-zip CSV, and on the skipped-existing and dry-run paths too.
 - Embedded project: **a live virtual layer no longer costs the recipient minutes per draw.** Such a layer re-runs its whole query for every feature count and every render — a canvas extent filter wraps the query rather than entering it — and QGIS pushes each equality in that query down as one filtered request *per outer row*. The packaged tables carried no index at all, so every one of those requests was a full scan of the inner table: in a delivered package, a 73,585 × 61,171 join took 7.4 minutes to show its feature count in the layer tree, and about as long again to draw. The columns a query compares with `=` are now indexed in each stratum's GeoPackage, turning each of those requests into a b-tree seek — the same package now counts in 13 seconds and draws in 18, for indexes built in hundredths of a second and about one percent of the file's size. This is the delivered-side half of the 0.3.0 fix, which removed the same nested-loop scan from the *packaging* by no longer asking such a layer for its extent. It makes a live virtual layer linear rather than cheap: the remaining per-row cost is QGIS's own, so a layer over millions of rows is still better materialized, or pushed into its source as a view and packaged as an ordinary layer.

 Version 0.4.0:
 - Virtual layers: **fixed a materialized virtual layer shipping empty in every package.** A virtual layer's definition can carry a `lazy` flag, which tells the provider to skip the query when the layer is opened — an authoring convenience, so that adding such a layer to a project does not immediately run its join against every source it touches. The flag rides along into the copy the packaging reads from, and a layer that has never run its query reports no fields and no features while still looking perfectly valid, so each stratum got a table with nothing in it. Nothing said so either: the layer was recorded as legitimately empty, and the check that reports features matching no stratum compares two numbers that are both zero. Such a layer is now loaded before its features are read, which is what the live route has always done by rebuilding the definition eagerly; the project file is untouched, and a layer left live is unaffected. Independently of `lazy`, a layer whose data provider exposes no fields at all now warns, naming the layer and its provider, instead of being written out empty in silence.
 - Processing: **`FULL_PACKAGE_PATH` now accepts an absolute path**, so the unpartitioned package can be published somewhere other than `OUTPUT_DIRECTORY` — a different drive or share than the per-stratum zips. It was the only one of the three path inputs still forced relative; `EXTRA_DIR` and `WARM_START_DIR` have always honored absolutes. A relative value behaves exactly as before, and an absolute one that happens to point inside `OUTPUT_DIRECTORY` is normalized back to the relative form, so it still bundles with a stratum zip of the same path instead of racing it. Only the basename of an absolute path is validated against the strict filename rules — the parent directories already exist and are the filesystem's business. Per-stratum zip paths (`ZIP_PATH_EXPRESSION`) are unchanged: they stay inside `OUTPUT_DIRECTORY`.

 Version 0.3.0:
 - Translations: **fixed the "Strata resolved" algorithm output staying English in every language.** The label was authored with `QT_TRANSLATE_NOOP`, which marks a string for extraction and hands back the source text — correct for a table built at import time, wrong here, because the outputs are declared from `initAlgorithm` and so are already past the point where the plugin translator is installed. Its three sibling outputs translate at that call and were unaffected. The Portuguese and Spanish translations existed and were finished all along; nothing ever looked them up, which is what makes this class of bug read as a missing translation. The output now translates like its siblings, and a test pins all four against a stubbed translator, since English output cannot tell a translated label from an untranslated one.
 - Embedded project: **fixed a live virtual layer costing minutes to hours per stratum.** Writing a project asks every layer for its extent, and a virtual layer answers by running its query over the stratum GeoPackage — where the packaged tables carry no index on the columns the query joins, making that a nested-loop scan whose cost grows faster than the stratum does. It was paid once per stratum, for a value the written project does not even store. Adding one such layer took a 93-stratum project from 35-55 seconds per stratum to 4 minutes on its smallest and 2 hours on a large one, with under 11 seconds of actual layer writing in each. The re-pointed layer is now left without a computed extent; the recipient's QGIS derives one on demand, locally, once. The delivered layer is otherwise unchanged — same query, same styling, and it still opens with its fields and features. Its subset string, until now silently dropped on the way into the package, is carried over too.
 - Processing: each stratum now also reports the **seconds spent building its embedded project** — the phase total, plus a breakdown naming every layer that build opens and the project write itself, slowest first. Only the layer writes were timed before, which is why a stall inside this phase survived two rounds of investigation. The two lines are read as a pair: when their totals agree the cost is in a layer the breakdown names, and when they diverge it is in the tree/style/relation work between them.
 - Embedded project: **remote layers are no longer re-fetched over the network once per stratum.** Layers riding only in the project (WMS/XYZ basemaps, annotations) were re-created for each stratum's project, and re-creating a layer re-constructs its data provider from the URI — which for a WMS source means a blocking GetCapabilities request. A project with 21 remote basemaps therefore made ~2000 identical requests across 93 strata, each one asking a server whose answer never changes. Those layers are now serialized once per run and rebuilt from that XML without ever opening their source, so the cost is independent of the network: against an unreachable host, a layer measured at 42 seconds per stratum now takes milliseconds. The packaged projects are unchanged — same sources, styles, names and tree positions — and still resolve their remote layers on the recipient's machine.
 - Processing: each stratum now closes with a **timing line** — total seconds spent writing its layers, plus the slowest few by name. Diagnosing a slow run from the log was not possible before: `qgis_process` block-buffers stdout, so a harness that timestamps the piped lines timestamps *flushes*, not work (dozens of per-layer lines land in one millisecond-wide burst), and a multi-hour stall inside a stratum could not be attributed to the layer that caused it. Carrying the measured numbers in the message body makes the attribution survive the buffering. One extra line per stratum; no behavior change.
 - Shared sources: **fixed a §12 group member whose subset the GeoPackage cannot evaluate still erroring once per stratum.** Grouping keys on the data source *ignoring* the subset — a subset is a per-member view over the same table, so it must not split the group — and the embedded project re-applies each member's subset to the shared table to tell the members apart. But a subset in the source provider's dialect (a PostgreSQL `lpad()`, a schema-qualified table SQLite has no schema for) can never run against a GeoPackage, so re-applying it failed inside GDAL (`ERROR 1: Failed to prepare SQL … no such table …`) one line per stratum, and the delivered layer showed no features. A member whose subset does not compile as SQLite is now **excluded from grouping**: it stages on its own, where the filter runs on the source provider and is materialized into its table, so the packaged project needs no runtime subset for it. Members whose subset the GeoPackage *can* run still share one table, unchanged. The previous entry gated re-apply to group members and fixed the *ungrouped* layer; this fixes the *grouped* one it left erroring.
 - Embedded project: **fixed a layer's subset string being re-applied to the packaged GeoPackage, breaking the delivered project.** A subset string is the *source* provider's SQL, but the embedded project re-applies it to the layer's GeoPackage table. A filter written for another provider — a PostgreSQL `::` cast, a schema-qualified table, a function SQLite lacks — is accepted by the layer API and only fails deeper, inside the data provider, where the failure surfaces as a bare `ERROR 1: Failed to prepare SQL …` on stderr and nothing else. The unusable filter was then stored in every stratum's project, so the delivered packages opened with that layer showing no features (the GeoPackage data was always correct). The subset is now re-applied only where it does work: a shared-source group (§12), whose one table holds every member's features and needs the subsets to tell them apart. An ungrouped layer's table already *is* its subset view, so nothing is filtered and nothing is demanded of the SQL. When a re-applied subset does not compile as SQLite the run now says so in plain language, naming the layer and the reason, instead of leaving a GDAL error as the only clue.
 - Embedded project: **fixed layers sharing a data source being left out of the packaged project.** Only one member of a shared-source group runs a write job, so the others have no result of their own; the project builder looked their outcome up directly and silently skipped every layer it did not find. Those layers were in the GeoPackage and in the report, but absent from the `.qgz`/stored project — the run gave no sign. They now inherit the group primary's outcome, as the reporting side already did.
 - Matching: **fixed the relation-chain memo never hitting.** It keyed on the whole chain, but packaged layers rarely share a chain outright — they share its leading hops and diverge at the final, layer-specific relation — so no two lookups ever matched and every layer re-derived hops another layer had just walked. On a 25-layer, 94-stratum project that was 4092 hop queries where 279 distinct answers existed (`0 hit(s), 2325 miss(es)` in the run log). Hops are now memoized per chain *prefix*, keyed also by the fields the next hop reads back, since one relation traversed for two different onward keys yields two different sets. Results are unchanged, and the run's `relation-chain hop memo:` line reports the saving.
 - Shared sources: a group whose members are **all** filtered no longer loses the primary member's own view. Clearing the primary's subset is what lets its table hold the union, but the plugin also forgot the subset, so that one member showed the whole union in the packaged project instead of its own features.
 - Processing: **fixed every staging and template step being logged twice.** `setProgressText` already writes its text to the algorithm's log — the Processing dialog appends the progress text to its log panel and `qgis_process` prints it to stdout — so the two steps that additionally pushed the same line as an info message emitted each line twice in the run log. The redundant push is gone; the progress label and the log line itself are unchanged.


