Restricted
This briefing is not public. Enter the passphrase you were given to continue.
Incorrect passphrase.
Marketing Intelligence · data lineage
How a campaign tactic’s timestamps travel from Workfront, Content Pantry and the DARF feed into the tables that answer “how long did this take?”. Flow runs left to right. Select any table to see its notes and trace what feeds it.
Analysis · Q1 2026 · revision _qmy_04092026
Q1 2026 told plainly — January to March 2026. You have just been handed this data and have never seen it before; nothing below assumes you have. Where a sentence needs a technical fact, the fact lives in an appendix and there is a pointer to it.
Tactics created between 2 Jan and 31 Mar 2026, across 103 campaigns.
Tactics that closed any phase in the quarter, whenever they began. 113 campaigns.
Just 12 of 540 tactics have both a targeting start and end date.
Tactics whose fulfillment “route” is the query giving up and substituting a start date.
source: analyses/q1_2026/story_qmy_04092026.md · reproduce: python tools/q1_story.py
Marriott marketing is organised like this:
signals_all_data_qmy_04092026.That is the whole apparatus. Somebody wanted to know how long a tactic takes and where the time goes, and since no single system knows the answer, the answer has to be assembled from crumbs. Nearly everything awkward in this document flows from that one fact.
There are twelve copies of that reporting table sitting side by side — dated drafts of the same query, kept as it was revised. They do not agree with each other. This document reads the copy the team’s own board labels “Most up to date query”. Which copy your report points at genuinely changes your answer.
You would think a quarter was a quarter. It is not. Imagine a bakery, and somebody asks how January went. You can answer three ways: the cakes we started in January; the cakes we finished in January (some of them were started in December); the cakes we planned to start in January (some of which were never made at all).
Three different sets of cakes, three different numbers, and all three are honest answers to a badly-worded question. The data has exactly the same three answers, and picking one silently is how two people end up with different figures for the same quarter.
| If you ask… | You mean | Tactics in Q1 2026 |
|---|---|---|
| STARTED | work we began in the quarter | 540 |
| WORKED | work we touched in the quarter, whenever it began | 565 |
| planned | what was scheduled to start in the quarter — a forecast | 480 |
It looks like the others and it is not the same kind of thing. It is derived from each tactic’s planned start date, so it answers “what was on the schedule” — a forecast, not an outcome. A report filtered that way is not measuring the quarter; it is measuring the plan for the quarter. Appendix A
So this document reports both real definitions, side by side, every time. That is not indecision. It is the only honest option, and after you have seen the two numbers next to each other a few times you stop finding it strange.
| Question it answers | Tactics | Campaigns | |
|---|---|---|---|
| STARTED | tactics we began in Q1 2026 | 540 | 103 |
| WORKED | tactics we did work on during Q1 2026 | 565 | 113 |
The oldest tactic in STARTED was created on 2 January 2026, because that is what STARTED means. The oldest in WORKED was created in June 2024 — a piece of work begun eighteen months earlier that only finished a phase now. Keep that picture in your head; it explains most of what follows.
Most of this work is email and display advertising, and most of it is authored in Workfront rather than Content Pantry. Appendix B
Here is the first thing worth telling anyone. A phase can only be measured when we have both a start date and an end date for it. Count how many tactics have that, phase by phase:
| Phase | Measured, STARTED | Measured, WORKED |
|---|---|---|
| intake | 181 of 540 — 34% | 197 of 565 — 35% |
| creative | 117 — 22% | 234 — 41% |
| targeting | 12 — 2% | 25 — 4% |
| fulfillment | 280 — 52% | 422 — 75% |
Twelve tactics out of five hundred and forty. Whatever you conclude about how long targeting takes, you are concluding it from a dozen rows.
This is not news to the people who built the system — it is written on their own board:
“WF Email subset is being mostly missed” · “non-cp email targeting = hard” Lucid board, verbatim
It is a known gap, not a fresh discovery. But a reader meeting a targeting average for the first time deserves to be told what it rests on.
Now the durations. Read the count before you read the average — an average over 12 rows and an average over 422 rows are not the same species of number.
| Phase | Typical (median) | Based on |
|---|---|---|
| intake | 36 days | 181 / 197 tactics |
| creative | 9 days (STARTED) · 39 days (WORKED) | 122 / 275 tactics |
| targeting | 14 days (STARTED) · 18 days (WORKED) | 12 / 25 tactics — thin |
| fulfillment | 40 days | 280 / 422 tactics |
End to end, from creation to fulfillment: STARTED runs a median of 40 days across 382 tactics; WORKED runs 56 days across 538.
Why is WORKED slower? Not because anything slowed down. It is the shape of the question. WORKED includes those long-running tactics begun in earlier quarters that only closed out now — of course they took longer; that is why they are still around. Statisticians call this a selection effect. You can call it “the survivors are the slow ones”.
And STARTED is biased the other way, which is the part people miss: a tactic still in flight has no end date yet, so it is simply absent. The slowest work of the quarter is missing from the “how fast are we today” number, by construction.
Neither number is wrong. Neither is the truth on its own. Report both, say which is which, and you will not mislead anybody. Means, ranges and the full per-phase table: Appendix C
None of this is an accusation. The system assembles an answer out of crumbs from three separate tools; these are the seams showing. But if you report from this table, you should know about all four.
The dates here are stored as text like 03/15/2026, not as real dates. A computer sorting
text puts 03/01/2026 before 12/01/2025, because 0 comes before
1 — it is alphabetising, not calendaring.
What it means for you: a date filter or a sort on this table can return a wrong answer that looks perfectly plausible. Anyone querying it has to convert first.
−447 dSix tactics in Q1 have a fulfillment end date earlier than their fulfillment start date. The worst “ended” in November 2024 and “started” in February 2026.
That is not a slow tactic; it is the start and the end being read from two different sources that do not share a clock. The table flags a handful of these itself (3 in STARTED, 5 in WORKED).
What it means for you: exclude negatives before averaging.
For each phase the table gives you a start date, an end date, and a duration. You would expect the duration to be the end minus the start. For intake, targeting and fulfillment it always is. For creative it is not — 2% of STARTED rows and 8% of WORKED rows disagree, because the dates and the duration are computed by two different pieces of the query.
What it means for you: if you recompute creative durations from the printed dates you will not reproduce the published number, and neither of you will be obviously wrong. Pick one, and say which you used.
99 tacticsThe table records which route each tactic took through the
business — and one of the most common fulfillment routes in the quarter is called
Tactic Start Date (Fallback).
That is not a route through the business. It is the query finding no real fulfillment signal and substituting the tactic’s start date so the row is not blank. It looks like data. It is a placeholder.
What it means for you: fulfillment coverage looks better than it is. Part of those 52% and 75% figures in section 4 is made of fallback.
Averages hide things; a single row does not. Two tactics were traced from the reporting table back into the three source systems underneath it, one from each group. Both check out — each reporting row matches the Workfront tasks recorded beneath it — and both illustrate the section 4 hole: neither has a real targeting date at all. Appendix H
Tactic Start Date (Fallback) placeholder rather than a real signal.Four honest gaps, so nobody has to find them the hard way:
Tactic Start Date (Fallback) route is meant to be this common, or has quietly become a crutch.creative_duration where it disagrees with its own dates.All four need the query definition read line by line, not more counting. That is
guides/3_SIGNALS_ALL_DATA_EXPLAINED.md.
Everything below is the same analysis with the technical detail left in. The document above is complete without it.
| Definition | Rule applied | Tactics |
|---|---|---|
| STARTED | tactic_creation_date falls in Jan–Mar 2026 | 540 |
| WORKED | any phase _end_date falls in Jan–Mar 2026 | 565 |
| planned | quarter = 1 AND year = 2026 (the view’s own columns) | 480 |
Overlap 261 · union 844 · WORKED-not-STARTED 304 · STARTED-not-WORKED 279.
The quarter, month and year columns are derived by the board from
ts.planned_start_date — “Add Quarters to the SQL Code… Based on
ts.planned_start_date::date”. That is why they answer a third question.
| Cohort | Question | Tactics | Campaigns | Creation dates span |
|---|---|---|---|---|
| STARTED | began in Q1 2026 | 540 | 103 | 2026-01-02 → 2026-03-31 |
| WORKED | worked on in Q1 2026 | 565 | 113 | 2024-06-03 → 2026-03-30 |
| tactic_type | created_platform | STARTED | WORKED |
|---|---|---|---|
| WF | 236 | 256 | |
| display media | WF | 226 | 246 |
| CP | 78 | 63 |
WF = authored in Workfront, CP = authored in Content Pantry. The test is one
email address: in tactic_signals, created_by = workfrontfusion@marriott.com means
WF; anything else means CP.
S = STARTED, W = WORKED.
| Phase | S n | S mean | S median | S range | W n | W mean | W median | W range |
|---|---|---|---|---|---|---|---|---|
| intake | 181 | 129.3 | 36.0 | 0 … 431 | 197 | 125.4 | 36.0 | 0 … 431 |
| creative | 122 | 18.6 | 8.5 | 0 … 70 | 275 | 55.4 | 39.0 | 0 … 603 |
| targeting | 12 | 22.2 | 13.5 | 5 … 77 | 25 | 60.8 | 18.0 | 0 … 463 |
| fulfillment | 280 | 57.1 | 40.0 | -60 … 310 | 422 | 56.4 | 41.5 | -447 … 313 |
End to end (tactic_creation_to_fulfillment_duration) — STARTED: 382
tactics, mean 52.6, median 40.0. WORKED: 538 tactics, mean 70.6, median 56.0.
Note how far the means sit above the medians on intake — a long tail of stalled intakes pulls the average from 36 days up to 129. The median is the more honest headline for every phase here.
The view’s own flags: is_bad_fulfillment_duration 3 (S) / 5 (W);
is_bad_activation_duration 2 (S) / 1 (W); tactic_launched 382 (S) / 538 (W).
All eleven *_date columns on signals_all_data_qmy_04092026 are typed
text and hold MM/DD/YYYY strings. Consequences:
where tactic_creation_date >= date '2026-01-01' errors outright — text compared to a date.ORDER BY sorts lexically, so 03/01/2026 sorts before 12/01/2025.to_date(nullif(col,''),'MM/DD/YYYY').The undated base revision signals_all_data types those same eleven columns as real
dates. This is a change the later revision made, and it is treated as a regression worth
reversing.
| Phase | STARTED rows < 0 | S min | WORKED rows < 0 | W min |
|---|---|---|---|---|
| intake | 0 | 0 | 0 | 0 |
| creative | 0 | 0 | 0 | 0 |
| targeting | 0 | 5 | 0 | 0 |
| fulfillment | 3 | -60 | 5 | -447 |
| tactic_id | Type | fulfillment_flow | Start | End | Days |
|---|---|---|---|---|---|
| 2972 | TLP Metadata | 02/20/2026 | 11/30/2024 | -447 | |
| 2973 | TLP Metadata | 02/09/2026 | 12/01/2024 | -435 | |
| 7203 | TLP Metadata | 03/23/2026 | 01/22/2026 | -60 | |
| 7476 | TLP Metadata | 01/29/2026 | 12/11/2025 | -49 | |
| 7175 | TLP Metadata | 02/23/2026 | 01/15/2026 | -39 | |
| 7902 | TLP Metadata | 03/12/2026 | 03/04/2026 | -8 |
All six are email and all six travel the TLP Metadata route. Across the whole table the
defect reaches -559 days and touches 36 rows on two routes.
| Phase | S comparable | S agree | S % | W comparable | W agree | W % |
|---|---|---|---|---|---|---|
| intake | 181 | 181 | 100% | 197 | 197 | 100% |
| creative | 117 | 115 | 98% | 234 | 215 | 92% |
| targeting | 12 | 12 | 100% | 25 | 25 | 100% |
| fulfillment | 280 | 280 | 100% | 422 | 422 | 100% |
Only creative disagrees, which usefully localises the problem: in that one phase the
_start_date/_end_date columns and the _duration column are produced by
different expressions.
| Route | STARTED | WORKED |
|---|---|---|
| (none) | 359 | 368 |
| External Agency | 181 | 197 |
| Route | STARTED | WORKED |
|---|---|---|
| (none) | 418 | 290 |
| Channel/Embedded Team | 122 | 275 |
| Route | STARTED | WORKED |
|---|---|---|
| (none) | 528 | 540 |
| CP DARF Enabled | 12 | 19 |
| WF Email | 0 | 6 |
| Route | STARTED | WORKED |
|---|---|---|
| TLP Metadata | 250 | 215 |
| (none) | 108 | 14 |
| Tactic Start Date (Fallback) | 99 | 95 |
| DME | 66 | 166 |
| CP Email DARF Enabled | 17 | 27 |
| WF Email | 0 | 48 |
(none) means the view recorded no route and no dates. Tactic Start Date (Fallback)
means it recorded a route and dates that are not real signals.
tactic_id = 7302| tactic_name | ACQ_LTO_BR_MTL_Bradesco_Q1_Email |
|---|---|
| campaign_name | 2026_ACQ_BR_Bradesco |
| tactic_type / platform | email / WF |
| fulfillment_flow | TLP Metadata |
| created | 01/14/2026 |
| creative | None → 03/19/2026 (64 d) |
| targeting | None → None |
| fulfillment | 01/14/2026 → 04/17/2026 (93 d) |
| creation → fulfillment | 93 days |
analytical_model.tactic_signals — status planned, created 2026-01-14,
last updated 2026-05-05, created_by = workfrontfusion@marriott.com → a WF
tactic.
| task_type | Tasks | First completed | Last completed |
|---|---|---|---|
| Tactic | 3 | 2026-04-20 | 2026-04-20 |
| Content | 2 | 2026-03-19 | 2026-03-19 |
| Language | 2 | ||
| Audience | 1 | 2026-03-19 | 2026-03-19 |
| NULL | 1 |
public.darf_tracking — no rows for this tactic, so its targeting and
fulfillment dates did not come from the DARF audience-delivery log.
tactic_id = 3828| tactic_name | 3828_ Q1_2025_ACB_Default_Email |
|---|---|
| campaign_name | ACB_Default |
| tactic_type / platform | email / WF |
| fulfillment_flow | WF Email |
| created | 01/17/2025 |
| creative | None → 01/22/2026 (370 d) |
| targeting | 01/17/2025 → 01/17/2025 |
| fulfillment | 01/17/2025 → 03/31/2025 (73 d) |
| creation → fulfillment | 73 days |
analytical_model.tactic_signals — status planned, created 2025-01-17,
last updated 2025-04-29, created_by = workfrontfusion@marriott.com → a WF
tactic.
| task_type | Tasks | First completed | Last completed |
|---|---|---|---|
| Content | 8 | 2026-01-22 | 2026-01-22 |
| NULL | 6 | 2025-02-10 | 2025-03-31 |
| Tactic | 3 | 2025-03-31 | 2026-01-22 |
| Audience | 1 | 2026-01-22 | 2026-01-22 |
| Targeting Intake | 1 | 2025-03-07 | 2025-03-07 |
| Targeting Production File | 1 |
public.darf_tracking — no rows for this tactic, same as above.
Note the shape of this one: created January 2025, creative phase closed January 2026 — 370 days. It is exactly the kind of row that makes WORKED look slower than STARTED, and exactly the kind of row STARTED cannot see.
Constraint first. These are production databases, so the walk was built to be cheap and auditable rather than exhaustive:
SET SESSION CHARACTERISTICS AS TRANSACTION READ ONLY and a 120s statement timeout.FILTER (WHERE ...), so answering the question twice costs almost nothing extra.LIMIT-ed. No row dumps.darf_tracking was never scanned. pg_index was read first, a btree on tactic_id found, and the table hit only with an equality lookup for a single traced tactic.Then the sequence.
pg_index for what is safe to filter on, pg_attribute for column types. That is what surfaced the first finding — every date column is text — before a single analytical query ran._duration = _end - _start, per phase. Prompted by an oddity in a traced row, it turned a one-off into a bounded, quantified defect.tactic_signals, signals_workfront_tasks and darf_tracking. This proves the lineage rather than assuming it.Two corrections the data forced, both worth knowing:
text column to a date and errored outright. The bug was the finding: these columns are MM/DD/YYYY strings.task_status on signals_placement_status_updates, copied from the board’s own signal card. The live column is status; task_status exists only on some dated variants. The board’s formula does not run as written against the table it names.Reproduce: python tools/q1_story.py.
Related reading: the High level tab (what the system is) ·
guides/3_SIGNALS_ALL_DATA_EXPLAINED.md (the query, line by line) ·
lineage/SOURCE_OF_TRUTH.md (exact columns and row counts).
All figures measured against heavy_processes, revision
signals_all_data_qmy_04092026.
Guide 1 of 3 · read this first
This is the map. Two kinds of statement appear below and they are kept apart on purpose: measured — read from the live databases; and the team says — written on the Lucid boards by the people who built this, quoted verbatim. That second kind is the why, and it exists nowhere else.
Workfront, Content Pantry, and a DARF file drop on Miacc SFTP.
staging-new is where work happens; heavy_processes is where work is measured.
A person exporting a spreadsheet — the step the team says is holding back automation.
Dated drafts of the same query, 3,385–4,245 rows, and they disagree.
source: guides/1_HIGH_LEVEL.md · measured 2026-08-26 · atlas: lineage/SOURCE_OF_TRUTH.md
Marriott marketing runs campaigns. Each campaign holds one or more tactics — a tactic is one thing a customer eventually sees: an email, a push notification, an SMS, or a DME (a display/digital piece).
Somebody asked a simple question:
How long does it take to get a tactic out the door, and where does the time go?
No single system can answer it. Work is planned in one tool, content is built in another, audiences are delivered by a third. So the answer has to be assembled from traces each system leaves behind. Those traces are called signals.
A signal is a timestamped fact that something finished — “creative intake was completed on 3 March”. Collect enough for one tactic and you can measure each stage.
The team is careful about what this is not. Their own banner:
“These are the Signals confirming completion — we don’t have them all yet.” board banner, verbatim
| Source | What it is | What it knows |
|---|---|---|
| Workfront (WF) | Project-management SaaS | Who did what, and when each task was actually completed |
| Content Pantry (CP) | The in-house authoring app | What the campaign and tactic are — names, types, content |
| Blob Storage (Miacc SFTP) | A file drop | DARF files — which audience/treatment rows were sent, and their status |
Three named jobs copy that into the analytics database. The team’s own automation table:
| Table | Automated? |
|---|---|
| signals_workfront_tasks | yes “Yes. Databricks Signals get data. Daily update.” |
| tactic_signals, campaign_signals | yes “Yes. Databricks Data_refresh_signals. Daily Update.” |
| darf_tracking | yes “Yes. Databricks DARF Tracking Refresh. Daily Update.” |
| signals_tactic_cp_wf | yes “Yes. Pulls from tables that are automated.” |
| signals_placement_status_updates_2 | no “No. Pulls from a workfront report on DME that needs to be refreshed and reuploaded to SQL.” |
There is also a fourth, quieter path the team mentions in passing: “Sharepoint to postgres automatically. Service Principle Account. Zapier as well.”
staging-newContent Pantry’s operational database. The live app writes here. ~1,300 tables.
heavy_processesThe analytics database. Copies of the above plus everything built on top. 260 objects.
Rule of thumb: staging-new is where work happens; heavy_processes is
where work is measured.
actual_completion_date.signals_all_data,
joins everything into one row per tactic with a column per milestone date. This is the
heart of the system.creative_start and
creative_end, the duration is subtraction.What follows is the model as the boards define it: five phases, each with an exact signal formula from the signal cards. It is the design intent, and it is worth knowing.
The deployed query does not implement it. signals_all_data_qmy_04092026
builds four phases — intake, creative, targeting, fulfillment. There is no In-Market
phase in the SQL, and no test on task_type = 'Audience' anywhere in it.
Targeting is derived from the DARF flags and a CP/WF platform test instead.
So this table tells you what the signals were meant to be. Where the board and the deployed SQL
disagree is catalogued in database/BOARD_VS_LIVE.md.
| # | Phase | Start / end signal | Where it comes from |
|---|---|---|---|
| 1 | Intake | cs.created_dt (campaign), ts.created_dt (tactic) | campaign_signals, tactic_signals |
| 2 | Creative | MAX(actual_completion_date) WHERE task_type = 'Creative Intake' | signals_workfront_tasks |
| 3 | Targeting | MIN(entry_date) WHERE task_type = 'Audience' → DARF T flag | signals_workfront_tasks, darf_tracking |
| 4 | Fulfillment | DARF P flag: MAX(file_name_date) WHERE status = 'P' | darf_tracking |
| 5 | In-Market | MAX(entry_date) WHERE task_status = 'DME Launched' | signals_placement_status_updates |
Two clarifications the team wrote down that you will otherwise get wrong:
“The ‘Final Status’ for placement tasks is DME Launched”board, verbatim
“In market status is end of targetting. For CP current day (when no darf)” board, verbatim — i.e. when a Content Pantry tactic has no DARF at all, targeting is closed out using the current date rather than a real signal
Email works differently from DME. The team spelled it out:
“creative starts at the creative intake task being completed through last content submissino task for WF EMAIL.”
“Fulfillment start for WF email is first content task completed.”
“Fulfillment end for WF email is last completion of a schedule for deployment task.” board, verbatim — misspellings as written
A tactic can be authored in Content Pantry or in Workfront, and the two follow different paths. The test is one email address:
“in tactic_signals created_by = workfrontfusion@marriott.com implies the tactic was made in WF. Otherwise it was CP”board, verbatim
Plus a rule of thumb from the team: “Anything with display media make platform WF.”
“sms = push for our tactic_type on tactic_signals and signals_workfront_tasks” board, verbatim — a trap
The milestone vocabulary inside signals_workfront_tasks:
Creative Intake Targeting Intake Targeting Production File Content Content Group Audience Placement Tactic Language Survey
T marks the end of targeting · P the end of fulfillment · D exists in the data — 71.7M rows — but appears in no diagram.
| Table | Rows | What it holds |
|---|---|---|
| public.signals_workfront_tasks | 343,527 | Every Workfront task with planned vs actual dates. The most important table here — nearly all timing comes from it. |
| analytical_model.campaign_signals | 4,493 | One row per campaign: name, type, owner, dates, status. |
| analytical_model.tactic_signals | 7,231 | One row per tactic: its campaign, type, dates, status, and the created_by that decides CP vs WF. |
| public.darf_tracking | 206,215,733 | The audience-delivery log — one row per attribute per treatment per file. The team warns: “Big table, queries will take around 2 minutes to load.” |
| public.v8_darf_files + v8_darf_tracking_v2 | 3,586 + 15,680,664 | The same DARF data restructured — file details factored out so they aren’t repeated on every row. 206M becomes 15.7M. |
| public.signals_placement_status_updates | 116,534 | When a DME placement changed status. The manual one. |
| public.signals_tactic_cp_wf | 8,626 | Tiny lookup: for each tactic, CP or WF. |
| public.signals_wf_email_fulfillment(_tasks) | 2,120 / 1,105 | Email delivery tasks. Being retired — see below. |
| View | What it does |
|---|---|
| workfront_tasks_clean | Trims the raw Workfront list to the tasks that matter |
| signals_tactic_cp_wf_2 | The CP/WF lookup, derived live rather than stored |
| signals_darf_status | Reduces the 206M-row DARF log to a status per tactic |
| Object | Rows | What it is |
|---|---|---|
| signals_all_data + 11 dated siblings | 3,385–4,245 | One row per tactic with every milestone date. The reporting table. |
| signals_no_darf | 15,989 | Same idea for tactics with no DARF involvement. |
| signals_debug, signals_debug_2 | small | Working views used while investigating. |
| analytical_model.CP_EXPORT_* | 1–360 | A separate small export path, unrelated to durations. |
signals_all_data, _qmy, _qmy_02112026, _03032026,
_04062026, _04062026_v2, _04092026, _04162026…
These are dated drafts of the same query, kept side by side as it was revised. Row counts
differ (3,385 to 4,245), so picking the wrong one silently changes your answer. Check which one your
report points at.
“Only thing not automated is the signals_placement Tasks”
“Needs to be ran each refresh, this is the step holding us back from automation.”
Someone exports a Workfront DME report to Excel and runs a Python script that drops and reloads the whole table. The team flags the consequence: “if someone deletes a placement task for some reason in WF it is in the void as far as our reporting is concerned.”
signals_wf_email_fulfillment_tasks is being retired — deliberatelyIts data looks stale, but that’s intentional:
“sweft is just the Schedule for Deployment Tasks. swt is already capturing this info and it is automated, so made the substitution.”
tactic_id is not the same type everywherebigint in most tables, character varying(255) in darf_tracking,
integer in some, double precision in the dated placement tables.
varchar-to-bigint fails loudly, which is fine. bigint-to-double-precision succeeds
silently — that’s the one to watch.
269,998 of 343,527 rows have task_type = NULL. The signal model only ever
sees the classified minority. The team’s note on nulls: “Not Automated: Any null values
likely imply tactic is new and has not been mapped yet.”
heavy_processes has zero triggers. Every table is filled by something
outside it — a Databricks job, the manual upload, or somebody running
REFRESH MATERIALIZED VIEW. If a job stops, nothing notices and no alert fires; the numbers
quietly stop moving.
staging-new is the opposite: 9 triggers, one of which is a
33,665-character function that rebuilds campaign_tactic_metadata whenever a
campaign or tactic is touched. That is the single largest transformation in the system, and it lives in the
operational database, not the analytical one.
Worth reading before you conclude you’ve found a bug — these are known:
“WF Email subset is being mostly missed.”
“non-cp email targeting = hard”
“audience_tactic_attribute does not track time it was added. tactic_signals does not track specific time it was added either.”so audience timing is partly unmeasurable today
“ACC should send us an audience count after it is processed. This is the end of Targeting”the intended future signal, not yet built
“Future: we nee to scrape the tactic level DARF for accuracy” and the same for treatment level
“Find out where nulls are and if we expect those nulls.”
Because nothing self-heals, ask how fresh the inputs are before trusting a number. Measured 2026-08-26:
| Table | Last activity | Age |
|---|---|---|
| signals_placement_status_updates | 2026-08-26 | current |
| signals_workfront_tasks | 2026-08-25 | current |
| campaign_signals | 2026-08-25 | current |
| signals_tactic_cp_wf | 2026-08-25 | current |
| v8_darf_files | 2026-08-25 | current |
| tactic_signals | 2026-05-19 | 99 days |
| darf_tracking | 2026-05-16 | 102 days |
| signals_placement_status_updates_2 | 2026-01-06 | 232 days |
The two ~100-day gaps have innocent explanations available — darf_tracking was
superseded by the v8_darf_* tables, which are current, and a quiet quarter with fewer
tactics launching looks exactly like this. Worth confirming with the team rather than assuming either way.
The habit is the point: check input dates before trusting an output number.
darf_tracking is 206 million rows and v8_darf_tracking_v2 is 15.7 million.
The team’s own warning — “Big table, queries will take around 2 minutes to
load.” Aggregate rather than scan, use LIMIT when inspecting values, and prefer the
pre-built artifacts over re-querying.
| I want to… | Go to |
|---|---|
| See the tables and how they connect, interactively | the Master view tab |
| Read a full worked analysis of one quarter | the Q1 2026 story tab |
| See what is wrong and what the options are | the Fix options tab |
| Look up exact columns and types | database/ddl/ |
| See what the diagrams claimed, for comparison | database/from_diagrams/erd_as_drawn.sql |
| Trace what feeds what, programmatically | lineage/lineage.json |
| Get the full validated reference | lineage/SOURCE_OF_TRUTH.md |
| Read what each original diagram said | diagrams/extraction/NN_*_extracted.md |
| Find the exact formula for one signal | diagrams/extraction/02_data_sot_current_day_extracted.md §2.2 |
| See SQL and screenshots that exist only as images | diagrams/extraction/00_image_only_components.md |
Source: guides/1_HIGH_LEVEL.md. Row counts and dates measured 2026-08-26
against heavy_processes and staging-new. Quotes with an amber rule are verbatim
from the Lucid boards, misspellings included — they are correct as written.
Assessment & options · written for a non-technical reader
No code, no jargon. One page of problem, two pages of options. Every figure is proved in the appendix, with a query you can run yourself.
Tactics with all four stages measured. That is 0.5% of the report.
Launch dates copied from a plan rather than observed — 1,895 rows.
Roughly one row in five. Not an engineering problem — a measurement problem.
Tactics with a real recorded launch the frozen snapshot cannot see.
source: maybe_fix.md · object under test: public.signals_all_data_qmy_04092026 · all figures measured 2026-08-26
The report is meant to tell us how long it takes to get a marketing tactic out the door. Today, about half the tactics are missing from it, and about half the answers that are in it are educated guesses presented as if they were measurements.
Every marketing tactic — an email, a display banner — goes through four stages:
There are 7,231 tactics in the system. The report contains 3,703 rows. Roughly half of all the work simply isn’t in the report. A1
Some of that is correct — cancelled work and genuine test records should be excluded. Some of it is accidental:
The report throws away any tactic whose name contains the word “test.” That catches real, live campaigns like “…App Benefits Banner TEST…”. A2
Real tactics dropped because their campaign name happens to contain a word on a block list (“LPA”, “100 days”, and three others). A2
The view’s predicate is tactic_type = 'display media' OR tactic_type = 'email'. No
push or SMS row can survive it. A2
Of the 3,703 rows that are in the report: A3
| Rows | Share | |
|---|---|---|
| Have all four stages measured | 17 | 0.5% |
| Have no stage measured at all | 1,249 | 34% |
| Launch date came from real evidence | 1,601 | 43% |
| Launch date was substituted from a plan | 1,895 | 51% |
| Nothing found, nothing substituted | 207 | 6% |
Only 17 tactics out of 3,703 have a complete four-stage picture. The report’s entire reason for existing — showing where time goes across four stages — is currently answered by half a percent of the data.
The “launched” column is worse than blank. When the system can’t find evidence that a tactic launched, it quietly substitutes the date it was planned to launch and marks it launched anyway. So a tactic that never shipped looks shipped, on a date nobody observed.
WF Email 917→917,
DME 568→568, CP Email DARF 118→116. _04092026 gained its 53
extra percentage points solely by adding the substitution arms. A4The team already knew. Their own note on the board:
“These are the Signals confirming completion — we don’t have them all yet.” board banner, verbatim · A13
The view behind the “average days per stage” charts computes one span per project and then reports that same number for Creative, for Targeting, and for Fulfillment. Different bars, same underlying number. A5
The live report is wired to frozen snapshots of its inputs, while fresher versions of those same tables sit right next to them, updating daily. A6
| The view joins | Frozen at | Live alternative | Last activity |
|---|---|---|---|
| signals_placement_status_updates_04062026 | 2026-04-06 | signals_placement_status_updates | 2026-08-26 |
| darf_tracking | 2026-05-16 | v8_darf_files / v8_darf_tracking_v2 | 2026-08-25 |
| analytical_model.tactic_signals (the spine) | 2026-05-19 | — | — |
| analytical_model.campaign_signals | 2026-08-25 | — | — |
| signals_workfront_tasks | 2026-08-25 | — | — |
For anything in the last four months, the report isn’t imprecise. It’s blind. And because it’s blind, those tactics fall through to the guessing logic — staleness manufactures fake numbers.
Eleven dated copies of the same report sit side by side. Sorting them by name gives you the wrong one — the two with the newest-looking names contain older logic. No single version has both the best logic and the freshest data — you pick one or the other. A8
This isn’t theoretical. The Excel workbook, the board charts, and the durations analysis each read a different copy. Two versions of the same activation chart were both shown to stakeholders: for 2024 display media, one says 142 days, the other says 57 days. Same population, same year, same board. A9
Three different kinds of problem are tangled together. They need three different kinds of fix, and only one of them is a coding problem.
Wrong table pointed at. Filters that catch real work. A column labelled “start date” that actually holds an end date — on 1,172 of 1,245 rows. A10 Negative durations — 36 of them — flowing straight into averages. A3 A “data quality” flag that marks unmeasurable rows as good, so a “clean” cohort is 56% padding. A11
These are real defects with real fixes, and they aren’t hard.
We took the 2,102 rows where the report is guessing or blank and asked: is the answer sitting somewhere else, just not wired up? A12
“Show me Q1” has three correct answers today: planned in Q1 (480 tactics), created in Q1 (540), worked on in Q1 (565) — sharing only 240 in common. A14 No fix makes that go away. Somebody has to decide which one the report means.
Same for the averages. The charts quietly apply different rules to different years — one year’s number cuts anything over 280 days as “human error,” the other doesn’t. That alone is enough to explain 142 days versus 57. A9
Underneath it, 79% of project-management records carry no task type and 78% carry no tactic ID at all A11 — so the report only ever sees a classified minority, and new work is the most likely to be unclassified.
The team said this themselves:
“We lose a lot of data with the given time constraints. In fact it keeps us from having anything at all on Targeting.”board, verbatim · A13
“audience_tactic_attribute does not track time it was added.”board, verbatim · A13
There’s also a structural quirk nobody can code around: in Content Pantry, editing an audience silently creates a brand-new audience, with no link back to the original. From a recorded call:
“If they originally started with audience 123, and they ended up deploying to audience 678, we don’t know at the time when they made the change… that’s the problem to solve.” Otter.ai transcript pasted on board 03 · A13
One step is a person exporting a spreadsheet by hand. The team’s own words:
“Needs to be ran each refresh, this is the step holding us back from automation.”
“As of now if someone deletes a placement task in WF it is in the void as far as our reporting is concerned.”lineage/notes.json · A13
When that spreadsheet doesn’t say which tactic a launch belongs to, the script grabs the
first four-digit number it finds in the task name. A task called “Q1 2026 Destination
Tiles” gets filed under tactic 2026 — which is a real tactic
(PREARRIVAL_RENAISSANCE_HOTELS-COPY). So launches don’t just go missing; some get
attached to the wrong tactic, silently. A15
There’s also a hand-maintained list of 124 placement names that decides what counts as a display tactic. Anything Marriott launches that isn’t on that list is invisible until somebody edits it by hand. A5
Not “every cell filled in.” That’s what got us here.
The report never guesses without saying so, never shows a number it can’t stand behind, and always tells you how much of the picture it actually has.
A smaller, honest report beats a full, fictional one.
The idea: the machine does everything. Nobody in the loop. The report gets smaller and far more trustworthy.
What changes:
A report that is right, refreshes itself, and needs nobody.
The headline numbers will drop, visibly. “Launched” goes from 92.5% to about 43%. Averages will move. Stakeholders will ask what happened. The honest answer is that the number was never real — but somebody has to say that out loud, once.
The 705 tactics with no record anywhere. They’ll show blank. That’s the point. Blank is the truthful answer.
The idea: everything in Option 1, plus a short weekly queue where a human supplies the handful of facts the machines genuinely don’t have. The report gets fuller and stays honest, because human input is labelled as human input.
Everything in Option 1, plus:
| Missing fact | Who already knows |
|---|---|
| Which display placements actually launched | Shelby (status list), Dalitso (the load) |
| Content Pantry process and edge cases | Ash, Natalie, Katie |
| Whether a tactic is DCA / MBOP-enabled | CJ |
| Audience revision history | Michele Grant, Marcus |
| Placement source of truth | Tytianna, Jared |
One person, about an hour a week, working a list the report writes for them.
A fuller report than Option 1, with every filled-in cell traceable to either a system or a named person. Coverage improves week over week instead of standing still.
An owner and a recurring hour. If nobody answers the queue, it degrades gracefully into Option 1 — still correct, just emptier. That’s the safety property that matters.
| Option 1 — Automated | Option 2 — Human in the loop | |
|---|---|---|
| Ongoing effort | none | ~1 hour/week, one owner |
| Report becomes | smaller, honest | fuller, honest, improving |
| Guessed numbers | removed | replaced with confirmed answers |
| Time to first version | weeks | weeks, then improves continuously |
| Fails how? | goes blank | degrades to Option 1 |
| Biggest risk | headline numbers visibly drop | nobody answers the queue |
These are not either/or. Option 1 is the foundation; Option 2 is Option 1 with a person attached. Build Option 1 either way.
The hand-run upload script carries a database password in plain text. Rotate it. A15
“ACC should send us an audience count after it is processed. This is the end of Targeting.”board, verbatim · A13
All queries are read-only against the heavy_processes analytics database. The repo ships a
guarded helper that opens the connection READ ONLY with a 120s timeout and refuses anything that
isn’t a SELECT:
PYTHONIOENCODING=utf-8 python tools/db.py "<query>"
Two cautions. public.darf_tracking is 206M rows and
public.v8_darf_tracking_v2 is 15.7M — always aggregate, never scan. And never use
information_schema on this database: it is filtered to the read-only role’s
privileges and silently under-reports (946 of 1,141 columns, 0 of 54 constraints). Use
pg_catalog.
The object under test throughout is
public.signals_all_data_qmy_04092026 — the revision the board flags as current.
All date columns in it are text in MM/DD/YYYY, so every query
below parses them with to_date(nullif(col,''),'MM/DD/YYYY'). Comparing them as strings
silently returns wrong rows. All figures measured 2026-08-26.
Claim: 7,231 tactics exist; 3,703 rows in the report.
SELECT (SELECT count(*) FROM analytical_model.tactic_signals) AS tactics,
(SELECT count(*) FROM public.signals_all_data_qmy_04092026) AS report_rows,
(SELECT count(DISTINCT tactic_id) FROM public.signals_all_data_qmy_04092026) AS distinct_tactics;
→ 7231 | 3703 | 3703
The row grain is clean — one row per tactic, no fan-out duplication. The loss is entirely the
WHERE clause, which is the last 12 lines of the view definition.
Claim: 151 production tactics dropped by the name filter; 326 by campaign-name filters; push/SMS excluded outright.
The view filters on ts.name NOT ILIKE '%test%' in addition to the explicit
is_test flag. Tactics that pass the flag but fail the name test, and follow production naming:
SELECT count(*) AS production_named_but_dropped
FROM analytical_model.tactic_signals
WHERE is_test = false AND name ILIKE '%test%'
AND (name ~ '^[0-9]{3,}_' OR name ~ '^(ACQ|RET|EMEA|CALA|APAC|MVW)_')
AND current_status <> 'cancelled'
AND tactic_type IN ('display media','email');
→ 151 — e.g. ACQ_Q4_LTO_US_Boundless_Chase_Review Res Animation TEST_FAIC_DME
SELECT count(*) AS dropped_by_campaign_name_filters
FROM analytical_model.tactic_signals ts
JOIN analytical_model.campaign_signals cs ON cs.id = ts.campaign_id
WHERE ts.is_test = false AND cs.is_test = false AND ts.current_status <> 'cancelled'
AND ts.tactic_type IN ('display media','email')
AND (cs.name ILIKE '%LPA%' OR cs.name ILIKE '%global module%' OR cs.name ILIKE '%100 days%'
OR cs.name ILIKE '%RETARGETING_MIGRATION%'
OR (cs.name ILIKE '%decision%' AND cs.name ILIKE '%engine%'));
→ 326
Push/SMS: the view’s predicate is
(ts.tactic_type = 'display media' OR ts.tactic_type = 'email'). No push or SMS row can survive it.
Claim: 17 rows complete; 1,249 wholly blank; 1,601 measured / 1,895 substituted / 207 nothing; 36 negative fulfillment durations.
SELECT count(*) AS rows,
count(*) FILTER (WHERE intake_duration IS NOT NULL AND creative_duration IS NOT NULL
AND targeting_duration IS NOT NULL AND fulfillment_duration IS NOT NULL) AS all_four,
count(*) FILTER (WHERE intake_duration IS NULL AND creative_duration IS NULL
AND targeting_duration IS NULL AND fulfillment_duration IS NULL) AS none_at_all,
count(intake_duration) AS intake, count(creative_duration) AS creative,
count(targeting_duration) AS targeting, count(fulfillment_duration) AS fulfillment,
count(*) FILTER (WHERE fulfillment_duration < 0) AS neg_fulfillment,
count(*) FILTER (WHERE creative_duration < 0) AS neg_creative
FROM public.signals_all_data_qmy_04092026;
→ 3703 | 17 | 1249 | 1232 | 809 | 454 | 1632 | 36 | 11
So each stage is known for only: intake 33%, creative 22%, targeting 12%, fulfillment 44%.
The measured/substituted split is readable directly off fulfillment_flow — the
view’s own label for which evidence it used:
SELECT fulfillment_flow, count(*)
FROM public.signals_all_data_qmy_04092026 GROUP BY 1 ORDER BY 2 DESC;
| fulfillment_flow | Rows | Meaning |
|---|---|---|
| Tactic Start Date (Fallback) | 1,245 | substituted — planned date copied in |
| WF Email | 917 | measured — a deployment task completed |
| TLP Metadata | 650 | substituted — planned start from a metadata table |
| DME | 568 | measured — placement reached DME Launch Ready/Launched |
| NULL | 207 | nothing found, nothing substituted |
| CP Email DARF Enabled | 116 | measured — a DARF file with status = 'P' |
→ measured 1,601 · substituted 1,895 · nothing 207
Observed-only cohort:
WHERE fulfillment_flow IN ('DME','CP Email DARF Enabled','WF Email')
SELECT '04092026' AS v, count(*) n, count(*) FILTER (WHERE tactic_launched) launched,
round(100.0*count(*) FILTER (WHERE tactic_launched)/count(*),1) pct
FROM public.signals_all_data_qmy_04092026
UNION ALL
SELECT '04162026', count(*), count(*) FILTER (WHERE tactic_launched),
round(100.0*count(*) FILTER (WHERE tactic_launched)/count(*),1)
FROM public.signals_all_data_qmy_04162026;
→ 04092026 | 3703 | 3424 | 92.5
→ 04162026 | 4031 | 1603 | 39.8
Why this is the whole argument. Compare the observed flow counts across the two:
WF Email 917 → 917, DME 568 → 568, CP Email DARF Enabled 118
→ 116. The evidence did not move. 04092026 gained its 53 extra percentage points solely by
adding the TLP Metadata and Tactic Start Date (Fallback) arms.
tactic_launched is defined in the view as fusi.fulfillment_end_date IS NOT NULL
— literally “the ladder produced a date from any arm, observed or imputed.” It is not a
launch status.
Claim: the view behind the per-stage duration charts computes a single project-wide span
and reuses it for every stage; the whitelist is 124 hand-typed entries; the SMS/Push channel is
dead code.
File: database/views/public.signals_no_darf.sql. Lines 182–192 —
the join is on project_id alone, with no restriction to the stage’s own tasks:
completion_dates AS (
SELECT cp.project_id, cp.project_name, cp.project_type, cp.id_campaign,
min(ct.actual_completion_date) AS first_completion,
max(ct.actual_completion_date) AS last_completion
FROM filtered_projects cp
JOIN signals_workfront_tasks ct ON cp.project_id::text = ct.project_id::text
WHERE ct.actual_completion_date IS NOT NULL
GROUP BY cp.project_id, cp.project_name, cp.project_type, cp.id_campaign
),
filtered_projects emits one row per (project_id, project_type). A project
classified as both Creative Email and Email Fulfillment therefore yields two rows carrying
identical first_completion / last_completion, computed over every
task in the project regardless of type. Creative, Targeting and Fulfillment all report the same number.
Two further defects in the same file:
WHEN project_type ILIKE '%Email%' … WHEN ILIKE '%Push%' … WHEN ILIKE '%SMS/Push%'. 'SMS/Push Fulfillment' contains Push, so arm 2 always wins and the SMS/Push channel never appears.analytical_model.tactic on id_campaign, so each project duration is averaged once per tactic in its campaign — weighting every project by how many tactics its campaign happens to hold.grep -o "'%[^']*%'::text" database/views/public.signals_no_darf.sql | sort -u | wc -l
→ 124
SELECT 'tactic_signals' t, max(created_dt)::date mx, count(*) n FROM analytical_model.tactic_signals
UNION ALL SELECT 'campaign_signals', max(created_dt)::date, count(*) FROM analytical_model.campaign_signals
UNION ALL SELECT 'signals_workfront_tasks', max(entry_date)::date, count(*) FROM public.signals_workfront_tasks
UNION ALL SELECT 'spsu (live)', max(entry_date)::date, count(*) FROM public.signals_placement_status_updates
UNION ALL SELECT 'spsu_04062026 (used)', max(entry_date)::date, count(*) FROM public.signals_placement_status_updates_04062026
UNION ALL SELECT 'v8_darf_files (live)', max(file_name_date)::date, count(*) FROM public.v8_darf_files
ORDER BY 1;
campaign_signals | 2026-08-25 | 4493
spsu (live) | 2026-08-26 | 118175
spsu_04062026 (used) | 2026-04-06 | 85657
signals_workfront_tasks | 2026-08-25 | 343527
tactic_signals | 2026-05-19 | 7231
v8_darf_files (live) | 2026-08-25 | 3586
SELECT max(file_name_date) FROM public.darf_tracking; -- → 2026-05-16 23:22:33
Why nothing self-corrects: heavy_processes contains zero
triggers. Every load is driven from outside — a Databricks job, the manual upload, or a
hand-typed REFRESH MATERIALIZED VIEW. If a job stops, no alert fires.
SELECT count(*) AS rows_after_cutoff,
count(DISTINCT tactic_id) AS tactics_after_cutoff,
count(DISTINCT tactic_id) FILTER (WHERE status IN ('LRY','LCD')) AS tactics_launched_after_cutoff
FROM public.signals_placement_status_updates
WHERE entry_date > DATE '2026-04-06';
→ 36896 | 668 | 467
(LRY / LCD are the live table’s status codes for DME Launch Ready /
DME Launched — the dated snapshots spell them out as names in a task_status column
instead. The two families are not interchangeable; see
database/BOARD_VS_LIVE.md §1.)
SELECT count(*) FROM analytical_model.campaign_signals WHERE created_dt > DATE '2026-05-19';
→ 149
SELECT c.relname, s.n_live_tup AS rows
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
LEFT JOIN pg_stat_user_tables s ON s.relid = c.oid
WHERE n.nspname = 'public' AND c.relname LIKE 'signals_all_data%'
ORDER BY c.relname;
| Object | Rows | Note |
|---|---|---|
| signals_all_data | 3385 | |
| signals_all_data_automated_refresh | 4245 | the only revision reading the live DARF path |
| signals_all_data_qmy | 3490 | |
| signals_all_data_qmy_02112026 | 3593 | |
| signals_all_data_qmy_03032026 | 3724 | |
| signals_all_data_qmy_04062026 | 3634 | |
| signals_all_data_qmy_04062026_v2 | 3634 | |
| signals_all_data_qmy_04092026 | 3703 | board says “most up to date” |
| signals_all_data_qmy_0416026_2 | 4245 | note the misspelt name (7 digits) |
| signals_all_data_qmy_04162026 | 4031 | later date, older SQL body |
| signals_all_data_v2 | 4245 |
Populations range 3,385–4,245 and are not nested — each contains tactics the others lack.
..._04162026 and ..._0416026_2 carry later dates but older SQL
bodies (they lack the TLP and fallback arms and use the superseded platform test) — someone re-ran an
earlier board panel on a later day. Sorting by name gives the wrong table.
Two more traps worth knowing before anyone writes a filter:
..._04162026 has no quarter / month / year columns at all — a report using them errors rather than returning a wrong number.signals_all_data_v2 computes identically-named quarter/month/year from fulfillment_end_date instead of planned_start_dt. Same names, opposite meaning.And the key structural fact: the only revision reading the live DARF path
(signals_all_data_automated_refresh) has none of the imputation arms, none of the campaign-name
exclusions, and a different DME status vocabulary. There is no revision that is both most-evolved
and current.
Both charts are pasted images on board 02, transcribed in
diagrams/extraction/00_image_only_components.md (items B8 and B9) — “ACTIVATION
DURATION (in Days, by Tactic In-Market Date)”:
| Version | 2024 Overall | 2024 DME | 2024 Email | 2025 Overall | 2025 DME | 2025 Email |
|---|---|---|---|---|---|---|
| B8 | 93 | 142 | 85 | 92 | 87 | 97 |
| B9 | 81 | 57 | 82 | 91 | 94 | 89 |
Three compounding causes, all documented:
database/board_sql/activation_avg_tactic_start_old_data.sql reads public.signals_all_data (3,385 rows); activation_avg_fulfillment_end_new_data.sql reads public.signals_all_data_qmy_03032026 (3,724 rows).AND fulfillment_flow <> 'Tactic Start Date (Fallback)'); the “old data” query has no such filter. Note that predicate is also NULL-rejecting, so it silently drops every NULL-flow row too — more than the analyst intended.“The drastic change in 2024 numbers is almost entirely due to the switch from tactic start date to fulfillment end date.”
“I believe someone deleted their placement tasks since those tactics still show up but not in our placement status reports anymore.”
“Hence, the activation duration is ‘short’ for DME in 2024 since the cutoff has this inherent bias.”
Related, in database/board_sql/duration_fulfillment_average_280day_cut.sql: a
WHERE "Campaign Days to Complete Project" < 280 cut applied to Fulfillment only, per the board
rule “fulfillments taking more than 280 days are treated as human error.” Removing the
upper tail from one stage and not the others biases every comparison downward.
The two arms sit in different CASE ladders in the same view and are gated on
different columns:
-- fulfillment_end_date, arm 5 (the VALUE)
WHEN ts_1.planned_end_dt::date < now()::date THEN ts_1.planned_end_dt::date
-- fulfillment_flow, arm 5 (the LABEL)
WHEN ts_1.planned_start_dt::date < now()::date
AND ts_1.planned_start_dt::date IS NOT NULL THEN 'Tactic Start Date (Fallback)'
SELECT count(*) AS labelled_fallback,
count(*) FILTER (WHERE to_date(nullif(r.fulfillment_end_date,''),'MM/DD/YYYY') = ts.planned_end_dt::date) AS end_eq_planned_end,
count(*) FILTER (WHERE to_date(nullif(r.fulfillment_end_date,''),'MM/DD/YYYY') = ts.planned_start_dt::date) AS end_eq_planned_start,
count(*) FILTER (WHERE r.fulfillment_end_date IS NULL) AS label_but_no_date
FROM public.signals_all_data_qmy_04092026 r
JOIN analytical_model.tactic_signals ts ON ts.id = r.tactic_id
WHERE r.fulfillment_flow = 'Tactic Start Date (Fallback)';
→ 1245 | 1172 | 383 | 72
The 72 are rows where planned start is past but planned end is still future: the label arm fires,
the value arm doesn’t. Consequence: never count launches with
fulfillment_flow IS NOT NULL — it over-counts by exactly 72. Use
tactic_launched.
This swap appears in no board panel — the direct ancestor ..._qmy_04062026
uses planned_start_dt in both places. It was made straight against the database.
SELECT count(*) FILTER (WHERE is_bad_fulfillment_duration = false) AS flagged_good,
count(*) FILTER (WHERE is_bad_fulfillment_duration = false
AND fulfillment_duration IS NULL) AS good_but_unmeasurable
FROM public.signals_all_data_qmy_04092026;
→ 3667 | 2071
The flag is CASE WHEN (end - start) < 0 THEN true ELSE false END. NULL < 0
is NULL, so every unmeasurable row falls to ELSE → false.
WHERE NOT is_bad_fulfillment_duration returns 3,667 rows of which 2,071 (56%) have
nothing to measure.
SELECT count(*) total,
count(*) FILTER (WHERE task_type IS NULL) null_task_type,
round(100.0*count(*) FILTER (WHERE task_type IS NULL)/count(*),1) pct_task_type,
count(*) FILTER (WHERE tactic_id IS NULL) null_tactic_id,
round(100.0*count(*) FILTER (WHERE tactic_id IS NULL)/count(*),1) pct_tactic_id
FROM public.signals_workfront_tasks;
→ 343527 | 269998 | 78.6 | 269365 | 78.4
Every stage signal in the report keys on
task_type IN ('Creative Intake','Content','Targeting Intake','Targeting Production File'), so it
only ever sees the classified 21%. Team’s note: “Not Automated: Any null values likely imply
tactic is new and has not been mapped yet” — i.e. the newest work is the most likely
to be unclassified, and therefore the most likely to be imputed.
WITH r AS (
SELECT tactic_id FROM public.signals_all_data_qmy_04092026
WHERE fulfillment_flow IN ('TLP Metadata','Tactic Start Date (Fallback)')
OR fulfillment_flow IS NULL),
w AS (SELECT tactic_id FROM public.signals_workfront_tasks GROUP BY tactic_id),
live AS (SELECT DISTINCT tactic_id::int tid FROM public.signals_placement_status_updates
WHERE status IN ('LRY','LCD') AND tactic_id IS NOT NULL),
p AS (SELECT DISTINCT tactic_id tid FROM public.v8_darf_tracking_v2 WHERE status = 'P')
SELECT count(*) AS guessed_or_blank,
count(*) FILTER (WHERE live.tid IS NOT NULL OR p.tid IS NOT NULL) AS recoverable_from_live_feeds,
count(*) FILTER (WHERE live.tid IS NULL AND p.tid IS NULL AND w.tactic_id IS NOT NULL) AS partial_wf_trace_only,
count(*) FILTER (WHERE live.tid IS NULL AND p.tid IS NULL AND w.tactic_id IS NULL) AS no_trace_anywhere
FROM r LEFT JOIN w ON w.tactic_id = r.tactic_id
LEFT JOIN live ON live.tid = r.tactic_id
LEFT JOIN p ON p.tid = r.tactic_id;
→ 2102 | 216 | 1181 | 705
This is the single most important query in the appendix. It is the evidence for the claim that this is not purely a SQL problem: 705 rows (19% of the whole report) have no launch evidence in any table — no placement record, no DARF file, no Workfront task. Re-pointing, re-joining and re-writing the SQL cannot produce a date for them.
The middle bucket (1,181) is the one worth funding: the trace exists in
signals_workfront_tasks, but no rule maps its task types to “finished.”
| Quote | File |
|---|---|
| “These are the Signals confirming completion — we don’t have them all yet.” | diagrams/extraction/02_data_sot_current_day_extracted.md (board header) |
| “We lose a lot of data with the given time constraints. In fact it keeps us from having anything at all on Targeting.” | diagrams/extraction/07_average_durations_extracted.md |
| “audience_tactic_attribute does not track time it was added.” | diagrams/extraction/03_signal_sources_wip_extracted.md @176365,19541 |
| “ACC should send us an audience count after it is processed. This is the end of Targeting.” | diagrams/extraction/02_data_sot_current_day_extracted.md |
| “Needs to be ran each refresh, this is the step holding us back from automation.” | lineage/notes.json → signals_placement_status_updates_2 |
| “…if someone deletes a placement task for some reason in WF it is in the void as far as our reporting is concerned.” | lineage/notes.json → same entry |
| audience-immutability exchange (Michele Grant / Marcus) | diagrams/extraction/00_image_only_components.md §D — Otter.ai transcript pasted on board 03 @26262,6847 |
| “The drastic change in 2024 numbers is almost entirely due to the switch…” | diagrams/extraction/02_data_sot_current_day_extracted.md (stakeholder review thread) |
The nine Lucid SVGs contain zero <text> elements — every
character was exported as an outlined glyph path and had to be decoded against a 373-glyph dictionary. All
nine now decode with zero unresolved glyphs, and quotes were render-verified. Two decoder defects were
found and repaired during extraction; both are documented in _archive/PIPELINE_LOG.md §C2 and
§C3. If a quote matters to a decision, open the render and confirm it.
WITH r AS (
SELECT tactic_id, quarter, year,
to_date(nullif(tactic_creation_date,''),'MM/DD/YYYY') AS created,
to_date(nullif(intake_end_date,''),'MM/DD/YYYY') AS a,
to_date(nullif(creative_end_date,''),'MM/DD/YYYY') AS b,
to_date(nullif(targeting_end_date,''),'MM/DD/YYYY') AS c,
to_date(nullif(fulfillment_end_date,''),'MM/DD/YYYY') AS d
FROM public.signals_all_data_qmy_04092026)
SELECT count(*) FILTER (WHERE quarter = 1 AND year = 2026) AS planned_q1,
count(*) FILTER (WHERE created >= DATE '2026-01-01' AND created < DATE '2026-04-01') AS created_q1,
count(*) FILTER (WHERE a >= DATE '2026-01-01' AND a < DATE '2026-04-01'
OR b >= DATE '2026-01-01' AND b < DATE '2026-04-01'
OR c >= DATE '2026-01-01' AND c < DATE '2026-04-01'
OR d >= DATE '2026-01-01' AND d < DATE '2026-04-01') AS worked_q1,
count(*) FILTER (WHERE quarter = 1 AND year = 2026
AND created >= DATE '2026-01-01' AND created < DATE '2026-04-01') AS planned_and_created
FROM r;
→ 480 | 540 | 565 | 240
quarter/month/year derive from ts.planned_start_dt
— when the tactic is planned to go to market. A campaign planned for March is often created the
previous October. So the quarter column is a forecast, and a report filtered on
it answers a planning question, not a delivery one.
Two edges: EXTRACT returns numeric, so year = 2026 works and
year = '2026' does not; and the text-date columns sort lexically, so
WHERE fulfillment_end_date >= '2026-01-01' compares '01/22/2026' >= '2026-01-01'
→ false for every row, silently.
File: database/board_sql/placement_status_etl.py (transcribed from board 02).
Line 54–55 — the comment and the regex disagree:
# Extract first 4-digit number at the start
match = re.search(r"\b(\d{4})\b", parent)
re.search with no anchor takes the first 4-digit run anywhere in the parent
task name. A parent named "Q1 2026 Destination Tiles" yields tactic_id = 2026, and
2026 is a real tactic:
SELECT id, name, tactic_type FROM analytical_model.tactic_signals WHERE id IN (2024,2025,2026);
→ 2026 | PREARRIVAL_RENAISSANCE_HOTELS-COPY | email
So placement events can be silently attributed to the wrong tactic.
Line 96 — if_exists="replace" drops and rebuilds the whole table each
run, which is why a placement task deleted in Workfront disappears from reporting with no trace.
Two more in the same file: pd.to_datetime(..., errors="coerce") turns an unparseable date into
NULL — “no signal” rather than an error; and float(...) on the id is why the
view joins spsu.tactic_id = ts_1.id::double precision, i.e. identifier equality through floating
point.
create_engine("postgresql+psycopg2://BTAdmin:<REDACTED>@main-processing…"). The
password is redacted in the repo copy; it is in plaintext on the Lucid board. Rotate it if
still in use.
Every name in the Option 2 table is drawn from the team’s own board notes, not inferred:
| Name | The note |
|---|---|
| Shelby | “Get list of statuses for DME from Shelby.” |
| Dalitso (Moyo) | “I am putting in the fixes to our two tables (Dalitso feel free to add any other changes made…)” |
| Ash, Natalie, Katie | “Reach out to Ash, Natalie, and Katie” / “CP = Ash, Natalie, Katie” |
| CJ | “CJ is the man. Let’s us identify which placement has what attributes to be DCA enabled or not.” |
| Michele Grant, Marcus | speakers in the audience-immutability transcript, diagrams/extraction/00_image_only_components.md §D |
| Tytianna, Jared (Brookins) | named as contacts for “Placement Source of Truth” |
| Katie R. | “Interim status check, Katie R doing this manually today, how can we automate this?” |
| What you want | Where |
|---|---|
| The deployed report’s SQL, reformatted and annotated | database/views/public.signals_all_data_qmy_04092026.sql |
| Line-by-line teaching walkthrough of that SQL | guides/2_REPORT_LOGIC_ONBOARDING.md |
| Revision inventory and which one to trust | guides/3_SIGNALS_ALL_DATA_EXPLAINED.md |
| Where the boards and the live database disagree | database/BOARD_VS_LIVE.md |
| Verified lineage — what feeds what | lineage/SOURCE_OF_TRUTH.md, lineage/lineage.json |
| Exact columns, types, keys, row counts | database/schema_snapshot.md |
| The team’s raw board content | diagrams/extraction/*_extracted.md |
| Content that exists only as pixels (charts, screenshots, transcript) | diagrams/extraction/00_image_only_components.md |
| How the boards were decoded, and every correction made | _archive/PIPELINE_LOG.md |
Source: maybe_fix.md. All measurements taken 2026-08-26 against
heavy_processes. Re-run any query above to reproduce.
Next steps · two asks for the Marriott team
Everything in the other tabs was measured against the database. These two asks are the things the database cannot tell us — they have to come from the people who own the report. Both are small to answer and both unblock a large amount of work.
Which copy is the report actually depending on — and which query in the DATA SoT model generated it?
Four aggregate numbers for one date range, from the team — not from the database.
Populations 3,385–4,245, not nested, and the newest name is not the newest logic.
Nothing to reconcile against. That is exactly what Ask 2 creates.
evidence for every number below: the Fix options tab, appendices A6 and A8
Right now there are multiple live copies. Please tell us the exact final live view you are depending on, and point us to the exact query in the DATA SoT model that was used to generate that copy.
Eleven dated copies of the same report sit side by side in public, all live, all queryable,
all disagreeing. Picking the wrong one silently changes every number a stakeholder sees.
The only revision reading the live DARF path (signals_all_data_automated_refresh)
has none of the imputation arms and a different DME status vocabulary. The revision the board calls current
(_qmy_04092026) reads snapshots frozen in April and May.
There is no revision that is both most-evolved and current — so
“which one are you using?” is not a filing question. It decides which set of trade-offs the
reported numbers already carry.
| # | What to send | Why it matters |
|---|---|---|
| 1 | The exact object name the live report/dashboard reads — schema and table, e.g. public.signals_all_data_qmy_04092026. | Fixes which of eleven populations every downstream number is drawn from. |
| 2 | Where that name is configured — the workbook connection, dashboard datasource, or scheduled job that points at it. | Tells us whether one place needs changing, or several disagree with each other. |
| 3 | The exact query in the DATA SoT model that generated that copy — the board panel or file, not a description of it. | Lets us diff the deployed view body against the intended query and list every difference. |
| 4 | Who refreshes it, and how often — person, job name, or “nobody, it’s a one-off.” | heavy_processes has zero triggers; if nobody refreshes it, the report is frozen and no alert fires. |
| 5 | Which of the other ten can be retired, or at least renamed with a _deprecated_ prefix. | Removes the single largest source of “why does your number differ from mine?” |
For a date range you choose, please give us a final aggregate report as your source of truth: what would the team say were the total intakes, total creative, total targeting and total fulfillment in that window — not read from the database. Aggregate numbers only, to keep it easy. Then have the internal team gut-check that those four numbers are 100% the source of truth. We will take it from there and dig into the tables to find exactly where the deviations are coming from.
Every problem in the Fix options tab was found by reading the database against itself. That method can prove a number is internally inconsistent. It cannot prove a number is wrong, because there is nothing outside the database to compare it to.
Four numbers the team stands behind changes that. It turns every discrepancy from an argument into arithmetic: your total minus the report’s total equals a gap, and every gap has a findable cause in a specific filter, join or fallback arm.
One small table. Pick any date range that is comfortable to answer for — a recent quarter is ideal, but any window the team can speak to with confidence works.
| Stage | Total the team stands behind | How you know / where it comes from |
|---|---|---|
| Total intakes requests that came in | ||
| Total creative creative completed | ||
| Total targeting audiences delivered | ||
| Total fulfillment tactics that went out |
Ask 1 fixes which numbers we are talking about. Ask 2 gives us something true to compare them against. Until both exist, every finding in this repository is a statement about the database’s internal consistency — which is useful, but is not the same as knowing whether the report is right.
Two things we’d like from you before we go further — both small.
1. Which live view are you actually depending on?
There are currently eleven live copies of signals_all_data in public, with row
counts from 3,385 to 4,245, and they are not nested — each one holds tactics the others don’t.
Could you confirm the exact final view the live report reads (schema and table name), where that name is
configured, and point us to the exact query in the DATA SoT model that was used to generate that copy? Also
useful: who refreshes it and how often, and which of the other copies we can treat as dead.
2. Your own aggregate totals, as a source of truth.
Pick any date range you’re comfortable with, and tell us what the team would say the totals were:
total intakes, total creative, total targeting, total fulfillment. Just four numbers — no row-level
detail needed. The important part is that they don’t come from the database, and that
the internal team has gut-checked them and is happy to call them 100% the source of truth. We’ll run
the same four totals against the report and dig into the tables to work out exactly where the deviations are
coming from.
Both asks are supported by measurements in the Fix options tab — appendix A8 for the eleven live copies, A6 and A7 for the frozen snapshots, A2 for the filters that drop real work, A3 for the measured/substituted split, and A12 for the 705 tactics with no record anywhere.