The Same New Evidence Produced Two Different Reviews
The review agent reopened an issue that an analyst had already resolved. It had the latest document and the current issue list, but not the reasoning behind that earlier decision. When we ran the same review with Hindsight memory enabled, the agent could retrieve that reasoning—and the issue stayed resolved.
That difference is the reason we built a replay into Chrimata. It runs the same incoming financial evidence through the due diligence agent twice: first without memory, then with Hindsight recall. Instead of trusting a “memory used” badge, we can inspect how the review changed.
Chrimata is a financial due diligence workspace that connects claims to source documents, tracks review issues, and records analyst decisions. Our question was simple: can an AI agent carry the reasons behind a previous decision into the next review without treating that decision as permanent?
Note — AI agent memory: ordinary application state can store that an issue is resolved. Agent memory is the context an AI system can retrieve later—such as what evidence supported the resolution, why the analyst accepted it, and what would make the conclusion stale.
A Resolved Status Does Not Explain the Decision
Imagine a company adds new evidence that changes monthly recurring revenue (MRR). One diligence issue related to that metric was already resolved. A review agent looking only at current numbers can notice the change but miss why a related question was closed. It may ask for the same evidence again or reopen work the analyst already completed.
The missing piece is not another copy of the status. It is the context behind the decision: the analyst’s explanation, the source documents, the date, and the conditions that would justify revisiting it.
In Chrimata, we record a decision receipt with those details. The receipt remains part of the application’s operational record, while Hindsight provides a memory layer that can bring relevant reasoning into a later review. PostgreSQL remains the source of truth for current issue and review state; Hindsight helps the agent recall context from earlier work.

Run the Review With Hindsight Off and On
The replay service takes a run, a Hindsight memory bank, and a trigger document. It calls the existing change-review pipeline twice. Both calls set persist=False, so a sandbox comparison does not create real review rows or open duplicate issues:
without = await self.change_review_service.create_ai_review(
db, run_id, bank_id,
trigger_document_id=trigger_document_id,
memory_enabled=False,
persist=False,
)
with_mem = await self.change_review_service.create_ai_review(
db, run_id, bank_id,
trigger_document_id=trigger_document_id,
memory_enabled=True,
persist=False,
)
The important part is that this uses the same evidence and application context on both sides. The only intentional change is whether the review retrieves prior context from Hindsight. The pipeline calculates metric changes and candidate issues before it builds the reasoning context, then conditionally adds recalled memories:
memory = []
memory_texts = []
if memory_enabled:
metrics = list({mc["metric"] for mc in comparison.metric_changes})
issue_types = [ci["metric"] for ci in candidate_issues]
memory = await self.memory_service.recall_for_change_review(
bank_id,
entity="Northstar Ops",
metrics=metrics,
issue_types=issue_types,
period=trigger_doc.document_date if trigger_doc else None,
)
for item in memory:
text = getattr(item, "text", "") or ""
if text:
memory_texts.append(text)
The retrieval query is specific to the review: it includes the entity, changed metrics, candidate issue types, and the new document’s period. That gives Hindsight useful context for recalling related decisions. The no-memory run skips recall, while both runs still use the current documents, claims, issues, and deterministic comparison results.
Compare Structured Outcomes in Python
We deliberately do not make a third model call to decide which review is better. The replay service compares structured issue IDs and outcomes returned by the two runs. For example, it checks which issues each side considers unaffected or reopened:
unaffected_without = {
ui.get("issue_id") for ui in without.get("unaffected_issues", [])
}
unaffected_with = {
ui.get("issue_id") for ui in with_mem.get("unaffected_issues", [])
}
reopened_without = {
ai.get("issue_id") for ai in without.get("affected_issues", [])
if ai.get("effect") == "reopened"
}
reopened_with = {
ai.get("issue_id") for ai in with_mem.get("affected_issues", [])
if ai.get("effect") == "reopened"
}
The server can then report a concrete difference: an issue was reopened without memory but remained resolved when Hindsight returned the prior decision. It also compares newly suggested issues and summaries, and reports how many memory texts were recalled.
That count is useful for inspection, but it is not a quality score. Eight retrieved memories do not automatically mean eight relevant memories. The reviewer still needs to look at the result and check whether the current evidence supports the decision.
What the Hindsight Memory Replay Showed
In our Northstar Ops walkthrough, the memory-off review reopened a related resolved issue. The memory-enabled run recalled eight prior memories and kept the issue resolved. The side-by-side result made the behavior visible: with Hindsight, the agent had access to the earlier rationale instead of treating the incoming evidence as a brand-new investigation.
That is a more useful demonstration of AI agent memory than showing a memory counter alone. We can inspect which outcome changed, which memories were returned, and whether the explanation still fits the latest documents.


Download the Architecture walkthrough as an MP4.
Hindsight Is Context, Not Authority
Memory should not freeze a past conclusion. If new evidence contradicts the analyst’s earlier reasoning, the agent should reopen the issue. Hindsight makes the old decision available; it does not decide whether that decision is still valid.
The replay is also not a controlled benchmark. The two sides use the same trigger document, but they are separate language-model calls. Model variation can change their outputs, and retrieved memories can be incomplete or irrelevant. The result shows what happened in a particular comparison; it does not prove memory always improves a review.
persist=False prevents the replay from committing its review and issue changes. The Hindsight path recalls memories but does not retain a new decision. The calls still consume inference time, and a later replay may differ. We treat it as a safe inspection tool, not a perfectly repeatable evaluation harness.
The split between systems matters too. The relational database tracks current operational state. Hindsight recalls prior reasoning. Current source documents remain the evidence the agent must evaluate. This boundary keeps memory helpful without turning it into an unreviewable source of truth.
What I Learned Building Agent Memory Replay
Make memory effects inspectable. Retrieval logs tell me that memory ran. A side-by-side replay tells me whether it changed an issue outcome or only the wording.
Keep sandbox runs from mutating the investigation. Both review calls use persist=False; candidate outcomes remain separate from committed analyst work.
Diff structured data with ordinary code. Issue IDs and state transitions are deterministic fields. Comparing them in Python is more predictable than asking an LLM to grade two prose summaries.
Let current evidence challenge the past. Hindsight can bring back why an issue was resolved, but the new document still determines whether that reasoning holds.
The replay changed how I think about AI agent memory. The hard question is not whether an agent can retrieve something from the past. It is whether that context affects a consequential next step—and whether a human can inspect why. Hindsight gives Chrimata a way to carry earlier reasoning into a new review; the replay makes that behavior visible before it becomes part of the investigation record.
The Chrimata source code on GitHub contains the review pipeline, Hindsight adapter, and replay implementation. For the memory layer itself, see the Hindsight GitHub repository, the Hindsight documentation, and Vectorize’s guide to AI agent memory.
Frequently Asked Questions
What is Hindsight AI memory?
Hindsight is a memory system for AI agents. In Chrimata, the review workflow stores prior decision context and retrieves relevant memories when a related financial due diligence issue appears again.
How does an AI agent memory replay work?
Chrimata sends the same trigger document through the change-review pipeline twice: once with Hindsight recall disabled and once with it enabled. It then compares the structured results server-side.
Does the replay change the live financial review?
No. Both replay calls set persist=False, so the candidate reviews are not committed as change-review records and they do not create new issue rows. The memory-enabled pass recalls context but does not retain a new decision.
Does Hindsight guarantee a more accurate due diligence review?
No. Memory can be incomplete or irrelevant, and separate model calls can vary. Hindsight provides prior context for the agent to evaluate; current source evidence still needs to support the conclusion.
Why use Hindsight instead of storing every decision only in PostgreSQL?
PostgreSQL remains the system of record for current ticket state and review records. Hindsight provides a way to retrieve semantically relevant prior reasoning when a later review needs context. The two systems serve different roles.