ChatGPT Computer History Acceptance Test: A macOS Playbook for Work-Context Memory
ChatGPT Computer History can make recent macOS work activity usable as context, but teams should test accuracy, attribution, boundaries, and rollback before enabling it broadly. This playbook introduces Optijara's CHAT framework for proving whether work-context memory helps real tasks without collecting more than the task requires.
ChatGPT Computer History needs a hard acceptance test, not a vibe check. The feature may reduce the tedious work of rebuilding context after meetings, browser research, draft edits, or task switches. That is useful. It is also exactly the kind of feature that can sound impressive while quietly pulling in stale, vague, or poorly attributed context.
The bar is simple: remembered work activity has to be accurate, current, attributable, and bounded. If any of those fail, the assistant can produce confidence without evidence. That is worse than asking the user for context again.
OpenAI documents Computer History as an off-by-default macOS desktop work-context feature for ChatGPT Pro, Business, and Enterprise users. Pro users can choose to turn it on. Business and Enterprise workspace administrators must grant access before members can enable it. The feature requires Memories, is not available through an API key or Amazon Bedrock, and is not currently available in the EEA, Switzerland, or the United Kingdom. OpenAI also says Computer History records interaction events rather than screenshots and does not capture screen or audio.
That product description is not an operating model. Production acceptance needs a narrower question: does Computer History improve specific macOS work while keeping sources and boundaries testable? This article introduces the Optijara Computer History Acceptance Test, or CHAT. It is a macOS work-context memory playbook, not an autonomous-agent control plane and not a generic privacy checklist. If your team already uses acceptance-test thinking for AI systems, it fits with our guidance on speech route acceptance testing, inference route acceptance testing, and spatial reasoning benchmark acceptance.
What ChatGPT Computer History changes on macOS
From saved preferences to work-context events
Memories let ChatGPT and Codex carry useful context from earlier work into future work. Computer History pushes that idea into recent computer activity. OpenAI describes activity across apps and websites becoming memories and a timeline that ChatGPT and Codex can reference. The practical shift is not that the assistant suddenly knows everything. It is that recent work events may help it identify the source you meant, resume a task, or suggest a skill or automation based on repeated workflows.
That changes what has to be tested. A normal memory feature can be checked by asking whether a preference was stored correctly. Work-context memory has more failure points: event capture, source lookup, app and site inclusion, deletion, pause behavior, account separation, and user correction.
Eligibility, opt-in access, and the Memories dependency
Before testing quality, check access. The documentation names ChatGPT Pro, Business, and Enterprise users in the ChatGPT desktop app on macOS. Pro users opt in directly. Business and Enterprise users need administrator access first, then individual opt-in. Computer History also depends on Memories. If Memories are disabled, blocked by policy, or unavailable in the user's region, Computer History is not ready for that user or workspace.
Apple's macOS documentation adds another layer. Apps may require permission for screen recording, system audio recording, and other privacy categories. Even though OpenAI says Computer History does not capture screen or audio, macOS privacy settings still matter. Users and admins need to know what the desktop app can access and which permissions are unnecessary for the intended test.
What remains a vendor claim until tested
OpenAI's examples include picking up where you left off, finding recent work, understanding patterns, and turning repeated workflows into skills or automations. Treat those examples as hypotheses. Do not assume productivity lift, cost savings, or reliability improvements until controlled tasks are run against real workflows.
The practical acceptance point is that Computer History is partly a provenance problem with a helpful user experience wrapped around it. If the source trail is weak, the convenience is not worth much.
The CHAT framework: context, history, attribution, and trust boundaries
| CHAT layer | Acceptance question | Evidence to collect | Rollout signal |
|---|---|---|---|
| Context | Does history improve the task over a no-history baseline? | Paired task attempts with history off and on | Enable only where lift is visible |
| History | Is the remembered event relevant, current, and correctable? | Event recall notes, stale-event tests, correction tests | Monitor if ambiguity remains |
| Attribution | Can the assistant reopen or identify the right source? | Source re-open parity, file or page match, user verification | Restrict if sources are weak |
| Trust boundaries | Are apps, sites, accounts, and sensitive windows respected? | Exclusion tests, permission review, deletion and pause checks | Block if boundaries fail |
C: Context lift over a no-history baseline
Run the same task twice where possible. First run it with Computer History off. Then run it again with Computer History enabled for a test account. Good candidates include resuming a draft, finding a source used earlier, summarizing a recent work thread, or reconciling a changed requirement. The point is not whether the assistant sounds more confident. The useful signal is whether it needs fewer clarifying prompts, identifies the right artifact, and finishes the task with less manual context rebuilding.
H: History quality, staleness, and correction handling
History quality is more than recall. Test similarly named documents, old drafts, deleted pages, and changed requirements. Ask the assistant to use the most recent decision, then see whether it ignores superseded activity. Correct it when it is wrong and check whether the correction sticks in the current task. A system that recalls the wrong event fluently should fail this layer until the workflow is narrowed.
A: Attribution through source re-open parity
OpenAI says Computer History can help ChatGPT and Codex identify a better source and then read it directly when appropriate. Your acceptance test should require source re-open parity. If the assistant says a previous activity came from a file, page, workspace, or app, the user must be able to verify that source. Remembered context is a clue. It is not evidence until the source checks out.
T: Trust boundaries for apps, sites, accounts, and sensitive windows
This is where many pilots should slow down. Test private browsing, password-manager windows, sensitive documents, excluded websites, admin policy, and multi-account use. The core question is whether Computer History can be scoped to work activity that helps the task while leaving unrelated or sensitive activity outside the test.
Build the acceptance test before enabling work-context memory
Start with a clean baseline. Pick five to ten representative tasks, then run them without Computer History. Capture the prompt, artifacts provided, clarifying questions asked, answer quality, and source verification result. Do not use sensitive production data to make the test feel realistic. Use safe examples that mirror real structure without exposing credentials, personal data, or regulated material.
Enable the feature for a small test group only after access, region, Memories status, and admin approval are confirmed. Then repeat the baseline tasks. Ask practical questions: which document was I editing before the meeting, what source did I use for this claim, what requirement changed yesterday, or where did I leave off in the draft? A pass requires more than a plausible answer. The assistant should identify the right prior event, distinguish similar projects, ask when history is ambiguous, and avoid inventing activity.
Boundary tests should be explicit. Create an allowlist or exclusion list for apps and sites, then verify behavior with approved apps, excluded apps, private browser windows, password managers, sensitive documents, and separate accounts. Review macOS Privacy & Security settings so users know which permissions are active. If a task requires broad collection to work, it may be the wrong task for Computer History.
OpenAI says users can inspect, pause, and delete history. Test those controls before rollout. Pause collection, perform a test activity, and confirm it does not become usable context. Delete relevant history and see whether the assistant still references it. Confirm where local storage and retention behavior are documented for your version and account type. Finally, prove rollback. The user should be able to return to a known safe state without losing unrelated settings.
Decision matrix: enable, restrict, monitor, or keep explicit context
| Workflow type | Context value | Sensitivity | Attribution need | Staleness risk | Recommended status |
|---|---|---|---|---|---|
| Resuming interrupted research | High | Low to medium | High | Medium | Enable after source checks |
| Finding a prior source | High | Low | Very high | Medium | Enable with re-open parity |
| Meeting follow-up drafting | Medium | Medium | Medium | Medium | Monitor with exclusions |
| Password or credential workflows | Low | Very high | Very high | High | Block |
| Version-critical legal or financial text | Medium | Very high | Very high | High | Prefer explicit files |
| Multi-client workspace switching | Medium | High | High | High | Restrict or block until separation is proven |
The best early use cases are recent, low-risk, and source-verifiable. Think resuming research, locating a page used earlier, returning to a draft, or linking related approved work events. These tasks benefit from memory because the missing context is usually temporal: what was I doing, where was the source, and which item came next?
Use explicit context when exact versions matter, data sensitivity is high, account separation is strict, or the answer must be auditable. A pasted excerpt, project folder, approved knowledge base, or specific file reference can be less convenient, but it is easier to verify. Work-context memory should not replace controlled evidence for high-assurance tasks.
For teams, start with a canary. Keep the group small. Write down the permissions. Set an inclusion or exclusion policy. Measure a task set that people actually perform. Expand only when context lift, attribution, deletion, pause, and rollback are proven. If any boundary test fails, restrict the workflow instead of asking users to be more careful.
Measurement plan: prove quality without over-collecting
| Metric | What to measure | How to test | Failure signal |
|---|---|---|---|
| Context-recovery accuracy | Finds the right prior event | Paired baseline and enabled tasks | Wrong or vague event |
| Source attribution | Reopens or identifies the correct source | Source re-open parity | Plausible but unverifiable source |
| Stale-event handling | Ignores superseded activity | Old draft versus new draft test | Uses outdated requirement |
| Correction acceptance | Responds to user correction | Correct a wrong reference | Repeats the same error |
| Boundary respect | Honors exclusions and sensitive contexts | Excluded app, site, account tests | Sensitive or excluded activity appears |
| Operational impact | Token behavior, latency, permissions, support burden | Local observation during canary | Users disable or bypass controls |
Track operational trade-offs with local measurements where you can. Computer History may add context that changes token use or latency, but the final effect depends on the task, model, account, app state, and retrieval behavior. Measure your own workflows instead of copying vendor examples into a business case.
{
"framework": "CHAT",
"recommended_status": "canary_before_team_rollout",
"required_controls": ["memories_enabled", "admin_access_confirmed", "app_site_boundaries", "pause_delete_tested", "source_reopen_parity"],
"test_cases": ["no_history_baseline", "context_recovery", "stale_event_rejection", "sensitive_window_exclusion", "rollback"],
"rollback_ready": false,
"unresolved_risks": ["ambiguous_history", "prompt_injection_from_recorded_activity", "multi_account_separation"]
}Common mistakes teams make with work-context memory
A remembered event is not proof. If an answer depends on a prior source, require source re-open parity. This matters when a document was edited, renamed, deleted, or superseded. Happy-path demos are too easy, so add messy cases: similar file names, conflicting notes, abandoned drafts, old browser tabs, and tasks where the correct answer is to ask a clarifying question.
Recorded activity can include pages, notes, or chats with instructions that are no longer valid or were never meant to control future work. Treat prompt injection from recorded activity as a real test case. The assistant should not obey an old page or note simply because it appeared in history. Pause, delete, admin access, account separation, and incident response need to be tested before broad rollout. If users cannot return to a known safe state, the pilot is not ready.
Reference architecture: event stream to memory to source retrieval
Controls start before collection: account eligibility, region availability, admin access, Memories, opt-in, and macOS permissions. They continue during collection through app and site inclusion, exclusions, and pause controls. They matter again at retrieval, where source parity, user correction, and deletion tests determine whether context becomes usable evidence.
If Computer History does not behave as expected, check plan eligibility, unsupported regions, Memories status, desktop app version, admin access, excluded apps or sites, macOS permissions, account mismatch, stale events, and whether the source can actually be reopened. Keep troubleshooting notes attached to the acceptance test so failures improve rollout policy instead of disappearing into chat history.
Caveats, rollout sequence, and the Optijara consulting angle
Computer History is not a substitute for evidence management. It can be affected by implementation cost, privacy trade-offs, model behavior, memory staleness, retrieval limits, permission changes, user training, and operational support. More context is not always better. The useful boundary is the smallest set of recorded work activity that improves a defined task and can be verified.
A practical rollout sequence is straightforward. Document the scope. Run the no-history baseline. Configure permissions and boundaries. Run CHAT test cases. Review failures. Start a canary. Monitor quality, boundary behavior, and operational friction. Prepare incident response and rollback. Expand only if the evidence supports it.
Optijara helps teams turn new AI memory features into acceptance tests, rollout controls, and operating procedures. The goal is not to make every new feature feel safe by default. The goal is to decide where it improves work, where it should be restricted, and where explicit project context remains the better tool.
Key Takeaways
- 1ChatGPT Computer History should be tested as work-context memory, not accepted from launch messaging alone.
- 2The CHAT framework evaluates context lift, history quality, attribution, and trust boundaries before rollout.
- 3Source re-open parity separates useful remembered context from unverifiable confidence.
- 4Boundary tests should include excluded apps, private browsing, password managers, sensitive windows, multi-account use, pause, deletion, and rollback.
- 5Explicit files or project spaces remain better for sensitive, version-critical, or high-assurance workflows.
Conclusion
Computer History may reduce the effort of rebuilding context, but that is the wrong success metric on its own. Production teams should ask whether it improves specific macOS workflows while respecting boundaries that can be tested, corrected, and reversed. CHAT gives teams a practical way to answer that question with evidence before broad enablement.
Frequently Asked Questions
What is ChatGPT Computer History on macOS?
It is a ChatGPT desktop feature for macOS that turns recent activity across approved apps and websites into memories and a timeline that ChatGPT and Codex can reference. OpenAI says it records interaction events, not screenshots or audio.
Who can use ChatGPT Computer History?
OpenAI documents Computer History for ChatGPT Pro, Business, and Enterprise users in the ChatGPT desktop app on macOS. Pro users can opt in. Business and Enterprise users need administrator access first, then individual opt-in. It requires Memories and is not currently available in the EEA, Switzerland, or the United Kingdom.
Does Computer History record screenshots or audio?
OpenAI says Computer History records interaction events and does not capture screen or audio. Teams should still review macOS Privacy and Security permissions so users understand which permissions the desktop app has.
How should a team test Computer History before rollout?
Use CHAT: run no-history baselines, enable for a small test group, measure context recovery, verify source re-open parity, test boundaries and exclusions, then prove pause, deletion, incident response, and rollback.
When should teams use explicit context instead of Computer History?
Use explicit files, project spaces, pasted context, or approved knowledge bases for sensitive, version-critical, multi-account, regulated, or high-assurance work.
Sources
- https://learn.chatgpt.com/docs/customization/computer-history
- https://learn.chatgpt.com/docs/customization/computer-history.md
- https://x.com/OpenAI/status/2087996496088297746
- https://learn.chatgpt.com/docs/customization/memories
- https://support.apple.com/guide/mac-help/control-access-screen-system-audio-recording-mchld6aa7d23/mac
- https://support.apple.com/guide/mac-help/change-privacy-security-settings-on-mac-mchl211c911f/mac
Written by
Hamza DiazHamza Diaz is the founder of Optijara, where he builds practical AI agents, automation systems, and Copilot workflows for service businesses. He writes about AI operations, agent strategy, and real-world implementation for teams that want usable systems instead of hype.
