# Understanding the data: a plain-language guide

This is the "what did we actually build, and why does it work this way"
companion. `docs/build-guide.md` is the technical build order; `docs/poc-domain-reference.md`
is the field-by-field reference. This document is neither — it's for building
the mental model, in plain language, with examples.

---

## 1. What are we actually storing?

TTC sells a subscription product about data-centre infrastructure: which
projects are being built, where, by whom, and how much capital is moving
into them. Underneath the four public-facing pages (Overview, Project
Pipeline, Market Monitor, Capital Tracker), there are **seven kinds of
records**:

| Table | In plain words | Real-world example |
|---|---|---|
| `companies` | Any organisation that plays a role in this world | "Equinix", "Blackstone" |
| `markets` | A place — a metro, or a country/region it rolls up into | "Northern Virginia", which rolls up into "United States" |
| `project_statuses` | The stages a project can be at, editor-configurable | "Announced", "Under construction" |
| `projects` | A physical data-centre project, tracked for its whole life | "Equinix's new campus in Ashburn, VA" |
| `project_status_changes` | An automatic log of when a project moved stage | "Moved from Announced to Financed on 2026-03-01" |
| `transactions` | A piece of capital activity — a deal | "Blackstone acquires a majority stake in Equinix's Ashburn campus" |
| `sources` | A citation backing up a fact | "Bloomberg article, https://..." |
| `power_policy_events` | A grid/permitting/policy signal for a market | "Dominion Energy approves new substation for Loudoun County" |

Everything else in the system — the paywall, the CRUD endpoints, the audit
trail — exists to manage and protect these seven things.

---

## 2. How the pieces connect

```
                     ┌──────────────┐
                     │   Company    │◄───────────────────────────┐
                     └──────┬───────┘                            │
                            │ operator (required)                │ buyer / seller /
                            │ owner, utility (optional)           │ target company
                            │                                     │ (all optional)
                     ┌──────▼───────┐        target_project      │
        Market ◄─────┤   Project    │◄─────────────────┐         │
        (required)   └──────┬───────┘                  │   ┌─────▼────────┐
           │                │  has many                │   │ Transaction  │
           │                ▼                          └───┤              │
           │        ProjectStatusChange                    └──────┬───────┘
           │        (auto-logged history)                         │
           │                                                      │
     ┌─────▼─────────────┐                                        │
     │ PowerPolicyEvent   │                                       │
     └────────────────────┘                                       │
                                                                    │
              ┌─────────────────────────────────────────────────────┘
              │        all four of: Project, Transaction,
              │        PowerPolicyEvent can cite...
              ▼
          ┌────────┐
          │ Source │  (a URL, a publisher — reusable across records)
          └────────┘
```

A few relationships are worth calling out because they're deliberate design
choices, not accidents:

- **A `Project` always has exactly one `Market` and one operator `Company`.**
  Not optional. You can't create a project floating in space — it has to
  belong somewhere and be run by somebody. It can *also* optionally have an
  owner company (if different from the operator) and a utility company.
- **A `Company` has no "role" field.** Equinix isn't stored as "type:
  operator" — it's an operator *because a project points at it*, and it
  could simultaneously be a buyer on a transaction. This stops the same
  company existing twice under two different roles.
- **A `Market` can point at another `Market` as its parent** (`parent_id`).
  This is how "Northern Virginia" rolls up into "United States" — same
  table, self-referencing, so market roll-ups don't need a second schema.
- **A `Transaction` can be about a company, a specific project, or both** —
  `target_company_id` and `target_project_id` are independent, not either/or.
- **A `Source` doesn't belong to any one record.** One Bloomberg article
  might back up facts on a project *and* a transaction *and* a market
  signal — so it's a many-to-many relationship (the `sourceables` table),
  not a foreign key sitting on `projects`.

---

## 3. The paywall: how "who sees what" actually works

This is the part that took the most care, because it's the whole reason the
product exists — TTC needs to actually withhold premium data, not just
visually hide it.

### The four tiers

```
Public  <  Registered  <  Subscriber  <  Internal
```

- **Public** — an anonymous visitor, or a failed/expired credential check.
- **Registered** — has a free WP account but no paid subscription.
- **Subscriber** — paying. Sees "unreleased" content (see below).
- **Internal** — TTC staff. Sees everything, including drafts nobody has
  published to anyone yet, and can create/edit/delete records.

The tier isn't something the caller (WordPress) gets to just claim — Laravel
double-checks it against WordPress on every request (see `docs/build-guide.md`
Part 1 if you want the mechanics). The point for this document: **by the
time any of the code below runs, the tier is a trustworthy fact**, not user
input.

### Where a tier actually comes from — this is WordPress's decision, not Laravel's

This is worth being precise about, because "Public / Registered / Subscriber
/ Internal" is **Laravel's own vocabulary** — it isn't how WordPress thinks
about users at all. Laravel doesn't decide who's a subscriber; it only ever
*asks* WordPress and trusts the answer. Two completely separate WordPress
systems feed into that one answer:

- **Subscription level** — TTC runs real **Paid Memberships Pro**, but tier
  resolution isn't raw PMPro; it goes through a bespoke mu-plugin,
  `ttc_membership_get_user_tier()`, which maps PMPro level IDs to TTC's own
  vocabulary: `none < free < pro < institutional`. This is what decides
  Public vs. Registered vs. Subscriber.
- **Staff access** — "Internal" isn't a subscription level at all. It's a
  WordPress **user role/capability** question (is this person logged in as
  an editor or administrator on the TTC site), which is a completely
  different WordPress concept from PMPro membership.

**The mapping between these two WP-native systems and this project's
four-value `Tier` enum is not finalized.** `docs/poc-domain-reference.md`
flags this explicitly: reconcile `none/free/pro/institutional` (and the
separate role/capability check for staff) against `Tier` when the actual WP
plugin gets built — don't assume `Public/Registered/Subscriber/Internal` is
the final vocabulary just because it's what the code uses today. Whoever
builds the WordPress side (Phase 1, Step 6 in `docs/build-guide.md`) decides
that mapping, not Laravel — Laravel's `ResolveEntitlement` middleware just
receives whatever tier WordPress's introspection response asserts (`tier:
"..."`) and binds it as-is. If WordPress's own role/membership logic changes
— a new PMPro level, a new staff role — nothing in this document's data
model changes; only the introspection response and the mapping do.

### Visibility is a clock, not a switch

Every gated record (`projects`, `companies`, `markets`, `transactions`,
`power_policy_events`) carries two timestamps:

```
subscriber_published_at   public_published_at
```

Think of it as a two-stage release:

```
   created            subscriber_published_at         public_published_at
      │                        │                              │
      ▼                        ▼                              ▼
  ──────────────────────────────────────────────────────────────────►  time
  (nobody sees it,      (subscribers + internal          (everybody sees
   except Internal)      see it now)                      it now)
```

A record is visible to a given tier the moment *now* passes that tier's
timestamp. Nobody flips a switch — visibility is computed fresh on every
single read, from the current time. That sounds like a small implementation
detail, but it's deliberate: the alternative (a stored `is_public = true/false`
flag, flipped by a scheduled job) fails silently the day that job stops
running. A record can go public early and stay that way, or never go public
at all, and nothing tells you it broke. Computing it from the clock every
time means there's no job to babysit and no way for the database to drift
out of sync with its own publishing schedule.

**A null `public_published_at` means "subscriber-only forever"** — some
content TTC may simply never make free.

### Two separate gates: which rows, and which fields

This is the one-sentence summary of the whole security model:

> Which **records** you can see, and which **fields** those records carry
> once you can see them, are two completely separate checks.

**Records** are filtered by a database query (`visibleTo($tier)`) — an
unentitled row is never even loaded into memory. **Fields** are filtered by
an allowlist in the API response layer — an unentitled column is never
written into the JSON.

The allowlist direction matters. Every resource in this codebase builds its
response as: "here are the fields *everyone* who can see this record gets,"
then "*if* you're subscriber-or-above, add these," then "*if* you're
Internal, add these." A brand new column added to any table is invisible to
every tier by default, until someone deliberately adds it to a list. That's
the opposite of the usual bug pattern (a denylist where a forgotten field
quietly becomes public).

Concretely, three buckets show up on almost every record type:

| Bucket | Who sees it | Examples |
|---|---|---|
| Base | Anyone who can see the record at all | name, location, dates, relationships |
| Premium | Subscriber and Internal | `capital_value_usd`, `power_price_mwh`, `vacancy_rate`, transaction values |
| Editorial | Internal only | `verification_status`, `internal_notes`, `last_verified_at` |

The "premium" bucket is specifically the *priced* data — the numbers TTC is
actually selling. Descriptive facts (a project's name, its market, its
timeline dates) aren't withheld just because a record exists at all; what's
gated is the financial intelligence layered on top.

### Sources are the one exception — visibility is borrowed, not owned

A citation (`sources`) has no publish timestamps of its own. If you can see
a project, you can see everything cited on that project — the citation's
visibility is *inherited* from whatever it's attached to, not independently
gated. There's no standalone "browse all citations" for anyone but Internal,
because outside the context of a specific record, a bare list of URLs isn't
paywalled content — it's just a research index.

---

## 4. How a project's life gets tracked automatically

A project isn't re-created every time its stage changes — "Announced" →
"Financed" → "Under construction" → "Operational" is **one row, updated in
place**. That's deliberate: it's what lets "how much capacity is under
construction in this market" stay a single, coherent number instead of
requiring you to de-duplicate project history first.

But if you only ever update the row, you lose the story of *when* things
changed. That's what `project_status_changes` is for — and you never have
to remember to write to it. A background observer watches every project
save:

```
Project created  ──────────────►  ProjectStatusChange: (from: nothing, to: Announced)
Project.status_id changed  ────►  ProjectStatusChange: (from: Announced, to: Financed)
Project.name changed only  ────►  (nothing logged — status didn't change)
```

This means the admin form, a future CSV import, and a future AI-assisted
entry tool all get a correct timeline for free, without each of them having
to remember to log it themselves.

`project_statuses` itself is a separate, editor-configurable table (an
editor can add "Paused" as a new status), but every status also carries a
fixed underlying **stage** — `pipeline`, `financed`, `under_construction`,
or `operational`. Dashboards always group by *stage*, never by the
editor-authored label, so adding "Paused" (mapped to `pipeline`) never
breaks an existing chart.

---

## 5. Who's allowed to change what

There's no login system inside Laravel at all — no `User` model, no
sessions, and critically, **no independent Laravel-side notion of "editor"
or "admin."** "Who is this" is resolved once per request from WordPress
(see §3's point about staff access being a WP role/capability check, not a
subscription level), and by the time it reaches any controller in this
codebase it's already been collapsed down to one of the four `Tier` values.
Given that:

- **Reading** most things (`projects`, `companies`, `markets`, `transactions`,
  `power_policy_events`, `project-statuses`) is open to every tier — the
  publish-window check (§3) is what actually limits what comes back.
- **Writing** anything — create, update, delete — currently requires
  whatever WordPress-side check resolves to `Tier::Internal`, checked
  everywhere, no exceptions. Right now that's a single undifferentiated
  bucket in Laravel's `Tier` enum. Whether WordPress itself distinguishes
  "journalist can draft, only an editor can publish" is a WordPress
  role/capability question that hasn't been decided yet — and if it is,
  the fix happens on the WordPress side of the mapping (§3) and in how many
  values `Tier` needs, not by inventing a separate permission system inside
  Laravel. (Flagged as an open question for TTC in `docs/build-guide.md`
  Part 10: "What are the internal roles — journalist, editor, admin — and
  what may each publish?")
- **`sources`** is the odd one out: even *reading* the standalone source
  library requires `Tier::Internal`, because outside a parent record
  there's nothing to inherit visibility from safely (§3).

---

## 6. The full shape of the API today

Every route below sits under `/api/v1/if/` and passes through the
entitlement check first.

| Resource | Read | Write |
|---|---|---|
| `projects` | everyone (tier-scoped) | Internal |
| `companies` | everyone (tier-scoped) | Internal |
| `markets` | everyone (tier-scoped) | Internal |
| `project-statuses` | everyone, unscoped (it's vocabulary) | Internal |
| `transactions` | everyone (tier-scoped) | Internal |
| `power-policy-events` | everyone (tier-scoped) | Internal |
| `sources` (standalone) | Internal only | Internal |
| `{project\|transaction\|power-policy-event}/{id}/sources` | inherits the parent's visibility | Internal (attach/detach) |

A record that fails its tier check doesn't come back as "403 forbidden" —
it comes back as **404 not found**. That's intentional: a subscriber-only
draft shouldn't even reveal that it exists to a public visitor.

---

## 7. Walking through one project, start to finish

1. **A journalist creates a project** (via the future WP admin, calling
   `POST /projects`) with a market, an operator company, and a status of
   "Announced". It's created with both publish timestamps `null` — a draft
   only Internal can see. `ProjectStatusChange` silently logs "created as
   Announced."
2. **It sits as an internal draft** while facts get verified. A `Source`
   (a press release URL) gets attached to it.
3. **An editor sets `subscriber_published_at` to now.** Subscribers can now
   see it — including its capital value, since that's a premium field —
   but the public still gets a 404.
4. **The deal financing closes.** The editor updates `status_id` to
   "Financed." `ProjectStatusChange` logs the transition automatically. A
   `Transaction` record gets created separately, linking the buyer/seller
   companies and, optionally, this project as its `target_project`.
5. **24 hours later** (per TTC's embargo policy — see `docs/build-guide.md`
   Part 10, still an open question with TTC on exact timing), an editor sets
   `public_published_at`. Now the public can see the project's name,
   location, and status — but still not its capital value, which stays
   subscriber-and-above forever unless that field's own gating changes.
6. **Anyone browsing the Market Monitor page** sees this project's capacity
   folded into their market's aggregate number — but a public visitor's
   aggregate and a subscriber's aggregate are *different totals*, because
   the aggregate itself is built from the same tier-scoped query as
   everything else (this part isn't built yet — it's Phase 3 in
   `docs/build-guide.md` — but the projects/companies/markets data
   this document describes is exactly what it will run on top of).

---

## 8. What's deliberately not built yet

So this doesn't read as more finished than it is:

- **Aggregates/dashboards** (Market Monitor, Capital Tracker pages) — the
  data model is ready for them; the query/endpoint layer isn't built.
- **`audit_logs`** — a generic before/after diff log for *any* field change
  on *any* record, with the actor who made it. `project_status_changes` is
  narrower (status only); this would be the broader version.
- **CSV import, AI-assisted extraction** — both explicitly "propose, never
  persist" by design (see `docs/poc-domain-reference.md`) — no code path
  exists yet from either into a published record.
- **The actual WordPress plugin** (token minting, introspection endpoint) —
  Laravel's side of the handshake exists and is tested; the WP side is a
  separate, not-yet-built codebase.
- **Editor vs Internal-admin distinction** — right now "Internal" is one
  undifferentiated tier for all writes.
