What AI Safety Means, and How to Judge Claims About AI Consciousness
October 5, 2026 5 min read
In practical terms, AI safety is not a quality a model has. It is the set of checks you put around a specific use of a tool: what the output will decide, who verifies it, and what happens when it is wrong. On consciousness, the honest answer is that the question is still open. The useful skill is knowing what would count as evidence, then holding every claim to that standard.
What does AI safety mean in practice?
The clearest public reference is NIST's 2024 Generative AI Risk Profile, a voluntary framework built around four activities: governing, mapping, measuring and managing AI risk. It names the risk categories that matter for generative systems, including data privacy, harmful bias, information integrity, information security and human-AI configuration, and it gives a working definition of confabulation: content that is false or erroneous but presented with confidence.
Two things follow from that framing. First, safety is not attached to the model alone. The profile holds that risk depends on the system, the use case, the context, the likelihood and the consequence, which means the same tool can be well handled in one setting and dangerous in another. Second, a framework is a mapping tool, not a certificate. Having walked one means you have thought about the risks; it does not mean the risks are zero.
For most readers the actionable layer is the operational one: the daily habit of checking outputs before they become decisions. That is where the next section goes.
How do I check an AI output before relying on it?
This is a proposed workflow rather than a measured standard, but each step has an observable pass condition:
- Write down what the output will be used for and what an error would cost. A drafting aid and a clinical summary sit at opposite ends of the same scale, and the bar should move with them.
- Extract every number, name, date and causal statement from the output. For each one, locate the source document and open it. A claim passes only when the document is in front of you and it actually says what the output said.
- Check scope. Look at what the source measured, on what sample and with what denominator, and respect what it did not test. An output that extends past the source's scope is an unsupported inference until you verify or qualify it; confabulation, in NIST's definition, is the stronger case — content that is actually false or erroneous but presented with confidence.
- Sort each sentence into three buckets: design (what was done), results (what was found), and not tested. Trust the first two within their stated scope; flag the third.
- Put a named person in charge of sign-off, and log what failed. The log makes the next review faster and turns individual vigilance into a team habit.
What would count as evidence for consciousness?
The current state of the question is genuinely unsettled, and the discomfort of that is worth sitting with rather than glossed over. A 2023 report by Butlin and colleagues, published on arXiv, took several competing theories of consciousness and derived computational indicator properties from them. Its assessment of the systems examined at the time suggested none of them were conscious, while finding no obvious technical barrier to implementing those indicators.
That report is a useful instrument and a limited one. It is theory-based and dated 2023, so it is not a proof about future systems, and passing an indicator would not by itself guarantee consciousness. What it does give is a vocabulary: specific, checkable properties you can ask a claim to address, instead of adjectives.
Vendor language is the other side of the coin. Product pages routinely describe models as aware, reflective or persistent, and those are first-party descriptions, not measurements. SentX's own homepage states that AGI and consciousness are research aims, not achieved results. That is a first-party stance, labeled as such, and it is the kind of scoping a reader should look for in any company's wording.
How should I assess a specific consciousness claim?
Five criteria, in rough order of usefulness:
- Who is making the claim, and what do they sell? First-party claims start discounted because the incentive structure is visible. Independent assessment carries more weight.
- Is there an operational test behind the word? If "conscious" reduces to "the system uses certain language or shows persistence," that is a description of output behavior, not a finding about inner experience.
- Which indicators were applied, and by whom? A self-assessment by the vendor is a statement of position, not an evaluation.
- Does the claim state its scope? Research aim and achieved result are different sentences. Hold the claimant to the weaker one unless the evidence forces the stronger.
- What would count against it? A claim that no observation could ever disconfirm is not a scientific claim; it is branding.
Where safety meets the claim: your data
Privacy is the one axis of AI safety a reader can act on today, and it is where the consciousness discussion has real teeth. If a system's memory shapes other users' interactions, what it remembers and shares is a safety property, not a feature footnote.
As a worked example from a published policy rather than a legal reading: SentX's privacy policy states that submissions are not confidential, that retained interaction representations can enter shared memory, and that submitted information can influence other users' interactions. It also states that account deletion does not promise removal of material already incorporated into shared memory or model weights, and the product's described memory is selective and changes over time, so recall is neither complete nor private.
The practical rule is to match sensitivity to policy. Contract terms, credentials, health details and anything proprietary stay out of tools whose policy routes input into shared systems. Drafting, summarizing and structuring public material are where the trade is usually worth making. And before assuming deletion cleans things up, check what the policy says deletion actually removes.
Both halves of the question resolve into the same discipline. Safety is the checks you run around a use, not a certificate the tool carries; consciousness is an evidence standard you demand of a claim, not a feeling you take from a product page. Find the source, read what it actually tested, and keep every claim inside that scope.