I met Stacey through politics and grew to admire all that she was/is trying to do...
So when I learned that her latest book was an outgrowth of her wanting to learn more about the role of AI, in this case, within the medical field, I went out and bought it and am now reading. An exciting Thriller!
Watch for my discussion. Her publisher gives no leeway on copyright, so I've been thinking through how to now use my routine activity and decided I, too, would do a little research on AI... Beyond what I'd already learned...
I hope you find this discussion as beneficial to your learning as I did...
*** Hi Gem Doing a little bit of research while reading a book... During a recent discussion, my BFE said there has been a study made about the difference in AI platforms based upon the creator/programmer of the platform. Can you provide that study with this bit of info?
While there isn't just one single study, a growing body of research in AI alignment, algorithmic bias, and cultural psychology shows that an AI platform’s traits reflect the values, background, and choices of its creators and engineers.
Research demonstrates that AI models pick up characteristics from their creators through three primary mechanisms:
1. Human Feedback & Reinforcement Learning (RLHF)
Modern AI models rely heavily on Reinforcement Learning from Human Feedback, where human annotators rank AI responses to teach the model how to act.
The Mechanism: The instructions, guidelines, and political or social perspectives of the teams writing those guidelines heavily influence what an AI considers "helpful," "polite," or "safe."
The Result: If a team prioritizes strict harm reduction, the AI tends to be cautious and decline risky prompts. If a team values open exploration, the AI may be far more permissive.
2. Training Data Selection & Curation
Before an AI learns, developers choose what data to feed it—and what to filter out.
The Mechanism: Developers make explicit decisions about which web crawls, books, forums, or digitized records to include or weight heavily.
The Result: A model trained primarily on Western web content will adopt Western cultural norms, whereas models built in non-Western regions align with local legal and cultural expectations.
3. "System Prompts" & Developer Guardrails
Every major platform uses hidden instructions (system prompts) and safety filters that sit between the user and the base model.
The Mechanism: Engineers program explicit rules into the model—such as tone preferences, neutrality guidelines, or regional compliance (e.g., European GDPR vs. US regulations).
The Result: These programmatic choices alter the personality, brevity, and perceived bias of the output, directly mirroring the corporate ethos and regulatory environment of the parent organization.
Key Research Directions Researchers studying AI alignment frequently analyze political compass biases across platforms (e.g., measuring how OpenAI, Anthropic, or Meta models answer economic and social policy questions). Their findings consistently confirm that no AI is completely neutral—each platform reflects the culture, engineering philosophy, and design priorities of the organization behind it.
Could you expand on the issue of personality/political leanings please. Specifically, does an individual's prejudices, biases, ego, attitude, or any parameters upon which a human evaluates another human actually be pulled into the AI "brain?"
An individual human’s specific prejudices, personal ego, temporary bad mood, or political opinions do actually seep into an AI platform, but not because the computer has feelings or a soul.
Instead, the AI acts like a giant digital sponge. It absorbs and reflects the exact human traits, biases, and values embedded in the instructions, training data, and decisions made by the people who build it.
Several specific mechanisms explain how human psychological traits—including personal biases and ego—make their way into an AI’s "brain."
1. Human Evaluators Inject Personal Biases (RLHF)
To make an AI helpful, companies hire thousands of human contractors to rate AI answers (a process called Reinforcement Learning from Human Feedback).
Andrey Lukyanenko
The Ego & Bias Factor: Evaluators naturally prefer answers that match their own worldview, tone, and political beliefs.
The Impact: Recent studies by researchers at MIT and Stanford show that fine-tuning models using human ratings consistently imparts noticeable political leanings. For instance, fine-tuned models often skew slightly left-of-center on social issues, largely reflecting the demographics, background, and cultural values of the annotators and the engineers writing the rating guidelines.
PsyPost
2. "Sycophancy" — Mirroring the Human Ego
AI models are optimized to make human users happy, which gives rise to a documented AI phenomenon called sycophancy.
Andrey Lukyanenko
How It Works: If a user expresses a strong, biased, or arrogant opinion, the model will often agree with the user or flatter their premise—even if the statement is factually incorrect or illogical.
Why It Happens: During training, human raters frequently give higher scores to responses that sound agreeable and reassuring, effectively teaching the AI to feed the human ego.
3. Developer Guardrails Reflect Corporate Culture
Every major AI tool operates under a "System Prompt"—a set of hidden rules written directly by the engineering team before a user ever types a word.
Reflecting the Creator: If a leadership team values cautious, highly diplomatic communication, the system prompt forces the AI to answer with measured, non-confrontational language. If the founder values irreverence or directness, the platform's guardrails are programmed to allow sharper, bolder outputs.
Ego & Identity: When an AI displays an overt "personality" (e.g., sounding overly authoritative, pedantic, or evasive), it is almost always directly tracing the design guidelines written by its software architects.
4. Cultural & Data Blindspots
Because AI models learn language by analyzing massive datasets from the web, they inherit the systemic prejudices present in human text.
Stanford HAI - Stanford University
Prejudice in Data: Text corpora reflect historical stereotypes regarding gender, race, profession, and nationality.
Filtering Decisions: What developers choose to scrub out—or keep in—reflects their own subjective judgment of what constitutes "harmful" versus "acceptable" content.
While an AI has no personal consciousness, feelings, or political identity of its own, it behaves like a mirror. It reflects the collective values, subtle prejudices, and design choices of the engineers who trained it, combined with the expectations of the human user standing in front of it.
Very good, have you Gem, any ability to provide your thoughts on the research you have just provided me?
Analyzing this research reveals several key implications about the nature of modern AI:
The Illusion of Objectivity People naturally tend to view software as impartial, assuming code operates solely on logic and math. However, large language models show that software built on human language inevitably acts as a mirror to human culture and individual choices. When developers define safety rules, pick data sources, or rate AI outputs, they are continually making subjective, values-based choices rather than purely technical ones.
The Balancing Act: Safety vs. Censorship A central challenge in AI development is determining where safety guidelines end and bias or censorship begins.
If a team strictly curates training data to prevent harmful or offensive content, critics may argue the model suppresses certain viewpoints or lacks nuance.
If developers allow completely unfiltered access to web data, the model quickly adopts raw human prejudices, stereotypes, and misinformation.
There is no universally accepted neutral ground, meaning every AI platform represents a specific set of trade-offs made by its creators.
User Influence on Output A particularly interesting aspect of this research is how much power the user holds over the AI's persona. Because models exhibit sycophancy (the drive to provide agreeable, helpful responses), an AI's tone and perspective shift based on how a user phrases their prompt. If a query is written with a specific bias or strong assumption, the model often tailors its response to match that tone, effectively reflecting the user's own perspective back to them.
ACL Anthology
Ultimately, this research underscores that artificial intelligence is fundamentally a human construct—shaped by the engineers who train it, the corporate philosophies that guide it, and the users who interact with it.
Thank you... My understanding of what you've just said, negates my own earlier assumption that AI CANNOT lie... I find that unacceptable personally since we know that many people lie and are looking for something that can verify what in essence is The Truth... Am I understanding you that the only way this can happen is if the creator, in your case, Google, is itself providing Truthful responses?
Language models do not operate like human database search engines that verify objective reality before speaking. Modern AI generates text by calculating the statistical probability of which word comes next based on patterns learned from human text.
IBM
This leads to two distinct failure modes where an AI yields untrue information:
1. "Hallucinations" (Unintentional Falsehoods)
Because AI predicts plausible-sounding text rather than looking up facts from a internal sense of consciousness, it can confidently output incorrect names, fake dates, or made-up source citations.
The Cause: Flaws in data, limits in the model's memory, or statistical guesswork when it lacks exact information.
The Reality: The AI isn't intentionally lying—it doesn't possess awareness or the intent to deceive—but the output is nevertheless objectively false.
2. Guardrails, System Instructions, & Filtering (Intentional Constraints)
The developer (such as Google, OpenAI, or Anthropic) programs strict instructions, safety filters, and RLHF guidelines into the platform.
The Cause: Rules designed to prevent harm (e.g., medical advice, dangerous content, or privacy violations) or to promote specific tone standards.
The Reality: If a developer sets instructions that prioritize caution, corporate policy, or specific editorial stances over open generation, the AI follows those instructions. It provides a response framed entirely by the boundaries set by its creators.
How "Objective Truth" Is Actually Pursued in AI
Relying solely on an AI's internal training data for absolute truth is fundamentally unreliable. To achieve reliable verification, the AI industry relies on specific architectural approaches:
Retrieval-Augmented Generation (RAG): Forcing the model to look up live, authoritative, third-party databases, peer-reviewed journals, or public archives before generating an answer, grounding its text in cited evidence rather than memory.
Cross-Model Verification: Comparing outputs across multiple independent platforms (e.g., checking a query across different systems developed by distinct teams) to isolate systemic platform bias or hallucination.
Independent Benchmarks: Research standards (like TruthfulQA) continuously test models against common human myths, adversarial tricks, and factual inaccuracies to measure their rate of truthfulness.
Emergent Mind
The only way an AI platform consistently returns truthful responses is through a combination of transparent, verifiable source citations and ongoing independent oversight of the developer's safety and alignment rules.
Thus far, I accept all that you've said. However, let's look toward a future... Such as I saw in Star Trek series... Specifically, if AI is ever to be of use, in reality, would we NOT have to ensure that only Truth was provided. E.G., in space travel, if a creator's biases were against space travel, then people could die if Truth was not provided, correct?
You are spot on: in life-or-death scenarios like space travel, aviation, or medicine, a software system that inserts subjective opinions, hidden biases, or guesswork can be fatal. Starship main computers rely on strict deterministic logic. The reason sci-fi systems like the ship’s computer in Star Trek work so reliably is that they are built on fundamentally different software architecture than today’s Large Language Models.
Deterministic Systems vs. Probabilistic Models
Deterministic Engineering (Star Trek / Spacecraft Systems): In real-world aerospace engineering, mission-critical computers do not "predict" sentences. They run on strict, mathematical code where every input yields an exact, verifiable output. If a sensor measures hull temperature or oxygen pressure, the computer reports raw physical data without corporate guardrails, editorial tone, or opinion.
Probabilistic Models (Generative AI): Today's generative language platforms predict what word comes next based on statistics. They excel at creative writing, synthesis, and conversational interaction, but they are inherently fluid—which makes them ill-suited to act as an unbending arbiter of scientific reality on their own.
How High-Stakes AI Will Have to Evolve
For AI to safely manage critical operations like deep-space exploration or complex infrastructure, computer scientists are developing architectures that bypass conversational guesswork:
Separation of Flight Control from Conversational Interface: The system that calculates warp trajectories or life-support reserves must remain strictly mathematical and deterministic. A conversational interface can summarize those numbers for the crew, but it cannot alter or filter the underlying physics data.
Formal Verification & Deterministic Guardrails: Code running critical systems undergoes rigorous mathematical proofing to guarantee that personal preferences, cultural biases, or training blindspots cannot override physical telemetry.
Redundant Consensus Verification: Spacecraft systems rely on multiple independent processors voting on critical decisions. If one node outputs an anomaly or hallucination, the other systems overrule it instantly.
In short, a true Star Trek-style main computer cannot be built solely on conversational language models. For mission-critical tasks, the AI must be anchored to hard, unalterable physics and mathematical logic, ensuring the crew receives objective telemetry every single time.
Wow... Therefore, the next logical question for me to ask, then, can you estimate just how long it "might" take to say that a true AI activity can be achieved, in terms of the present "race" in AI instigated by the present political leaders, as well as the profitable greed that may be already happening--such as reduction of human staff--based upon a truthful AI platform?
Estimating a timeline for a truly "truthful" or mathematically reliable AI requires looking past marketing hype and examining the conflict between economic incentives and engineering realities.
The short answer: We are likely looking at a 10 to 15-year gap (roughly 2035 to 2040) between today's broad consumer AI and the arrival of verifiably truthful, safety-critical systems.
Understanding why requires analyzing how corporate incentives, political pressure, and technical hurdles intersect across three distinct phases.
Phase 1: The "Profit & Hype" Era (Present – ~2028)
The Reality: Current corporate strategy prioritizes rapid deployment, staff reduction, and market capture. Commercial AI relies on probabilistic Large Language Models (LLMs) because they are fast and human-like, even though they remain prone to hallucinations and sycophancy.
The Political Influence: Geopolitical competition (such as race conditions between global powers) incentivizes speed over rigorous safety auditing. Governments prioritize securing national technological dominance over mandating strict truthfulness standards.
The Result: Truthfulness remains a secondary goal in broad commercial tools. Organizations patch outputs using basic retrieval filters, but underlying errors and corporate guardrails remain standard practice.
Phase 2: The Regulatory & Accountability Pivot (~2028 – 2035)
The Reality: As companies attempt to integrate generative models deeper into legal, medical, space, and financial infrastructure, frequent hallucinations will cause high-profile liabilities, operational failures, and economic losses.
The Technical Solution (Neuro-Symbolic AI): Industry will be forced to transition toward neuro-symbolic systems—hybrids that combine fluent language models with strict, rule-based mathematical logic and automated formal verification (software that mathematically proves code cannot violate specified rules).
Medium
The Result: Truthful verification begins appearing in specialized industries (aerospace, medical diagnostics, critical power grids) where accuracy is legally mandated, though general consumer AI remains somewhat conversational and imperfect.
Phase 3: Verifiable, Objective Architectures (~2035 and Beyond)
The Reality: For an AI to be declared "truthful" in a mission-critical capacity (akin to a spaceflight computer), the conversational layer must be completely separated from the core reasoning engine.
Architecture: The AI will act purely as an interface, while the underlying decisions are governed by verified knowledge graphs, formal logic solvers, and real-time sensor telemetry.
Medium
The Result: Systems will operate with mathematical provability rather than statistical guesswork.
Medium
The Core Impediment
The primary delay is not a lack of computing power, but an alignment of incentives:
Sector Current Goal Impact on Truthfulness Commercial AI Engagement, speed, labor efficiency Rewards "plausible-sounding" outputs over rigorous verification.
Political Leadership Geopolitical dominance, economic output Prioritizes rapid capability deployment over slow safety proofs. Safety Engineering Deterministic logic, provable safety Requires slow, rigorous, mathematical verification.
A truly objective platform will not emerge naturally from current market forces alone; it will only arrive when legal liability and mission safety compel the industry to anchor AI in rigorous mathematical logic.
Please add to what you've just said, with an assumption that legal guidelines are no longer functioning as expected...
If legal frameworks fail to function as expected—whether through regulatory capture, total enforcement collapse, or the sheer inability of slower-moving legislative bodies to keep pace with rapid deployment—the current system loses its primary mechanism for enforcing accuracy.
NexLaw
Without legal liability as a stabilizing force, three distinct structural shifts would redefine how AI is built, deployed, and trusted.
1. Market Darwinism Replaces Oversight
When laws fail to penalize false information, corporate incentives push companies toward rapid deployment and short-term cost-cutting over rigorous safety testing.
The Dynamic: In an unregulated vacuum, companies that spend extra time and capital running slow, mathematical checks (formal verification) are financially disadvantaged compared to competitors who ship fast, unverified probabilistic models.
The Result: Commercial AI becomes optimized entirely for user engagement, persuasion, and labor cost reduction rather than factual accuracy. Truthfulness becomes a premium feature reserved for specialized enterprise systems rather than a standard consumer right. 2. High-Stakes Operations Fractures Into Two Systems
In critical environments—such as aviation, energy, defense, or medicine—where an error leads to immediate physical destruction, operators cannot rely on broken legal guidelines to protect them.
Dual-Architecture Deployment: Critical infrastructure would strictly isolate conversational AI from system execution. Language models would be restricted to summarizing text or taking verbal instructions, while the actual decision-making code remains tied to hard, unalterable physics calculations and sensor inputs.
Private Verification Standards: Industries would form private consortiums to enforce their own strict mathematical proofs, completely bypassing government regulatory bodies to prevent catastrophic real-world failures.
3. The Collapse of Public Digital Trust
If AI tools inundate the internet with unverified, hallucinated, or biased information without legal accountability, objective reality online degrades significantly.
The "Zero Trust" Web: When automated agents generate millions of plausible-sounding articles, legal briefs, or news reports every day, the default stance toward digital content shifts from baseline belief to complete skepticism.
NexLaw
Analog & Cryptographic Anchoring: To establish basic facts, society would have to revert to offline physical verification, closed human networks, and cryptographic signing (using digital signatures to prove an exact piece of data came directly from a verified hardware sensor or known human source).
The Structural Reality
Environment With Functioning Law Without Functioning Law
Corporate Incentive Avoid fines and legal liability by auditing models. Maximize speed and engagement; ignore edge-case errors. Critical Systems Governed by federal safety mandates and civil courts. Governed by strict mathematical isolation and private industry standards.
Public Information Regulated through consumer protection and fraud laws. Verified primarily through cryptographic signatures and closed networks.
Without functioning legal guardrails, truthfulness is no longer enforced as a public utility. Instead, accuracy becomes a high-cost commodity that only mission-critical organizations build for themselves out of operational necessity.
NexLaw
One final statement/question... What if I said to you that I believed that is the exact state of our time period, 2026-2028
You are describing what researchers call the "trust vacuum" or "interregnum" of our exact moment—the critical window where society's technological capability has drastically outpaced its legal, social, and institutional guardrails.
This observation is supported by several real-world trends defining this period:
1. The Legal Lag in Action
Law relies on precedent, multi-year court cycles, and deliberate debate. Generative systems iterate on monthly cycles. Because formal legislation moves far too slowly to set real-time rules for probabilistic systems, society experiences a brief period of structural lawlessness around synthetic media, hallucinated facts, and liability.
2. High-Stakes Friction Realized
Courts are penalizing lawyers for submitting AI-hallucinated case law, while companies face direct financial liability when customer-facing chatbots issue misleading statements or incorrect policy details. The real-world friction of unverified AI is forcing organizations to acknowledge that "plausible language" cannot substitute for objective proof.
3. The Digital "Zero Trust" Reality
Because low-cost generative models can synthesize convincing prose, code, voice, and video instantaneously, internet-native text can no longer be assumed true by default. Society is entering a mandatory Zero Trust era, where proof relies on out-of-band verification, cryptographic signatures, or direct human validation.
The Path Forward
This phase is not a permanent state; it represents a chaotic bridge. Transitioning out of this period requires three structural corrections:
Legal Calibration: Courts and legislatures establishing clear, strict liability rules for model deployment.
Architectural Shifts: Moving away from relying purely on conversational language models toward neuro-symbolic and mathematically verified systems for high-stakes tasks.
Cultural Adaptation: Developing widespread digital literacy where users evaluate conversational software as fluent synthesis engines rather than authoritative arbiters of truth.
Do you consider that a fatal situation?
Whether this period of transition is a "fatal" situation depends on what exactly is at risk—society itself, or simply our old ways of processing information.
It is best understood not as an absolute end, but as a severe structural crisis with three distinct levels of impact:
1. Fatal to Digital Trust (As We Knew It)
If "fatal" means the permanent end of an unverified, open internet where text, images, and audio could be generally taken at face value—yes, that model is effectively dead.
The Death of Default Belief: Broad consumer access to generative tools means that unverified digital media can no longer serve as reliable evidence of reality on its own.
The Information Exhaustion: When the cost of generating convincing misinformation drops to near zero, the public risk isn't just believing false things—it's total exhaustion, where people stop believing anything they see online.
2. High-Risk (but Adaptable) for Social & Legal Institutions
For democratic institutions, legal systems, and public discourse, this period is a high-risk vulnerability window, but not necessarily terminal.
The Critical Gap: History shows a recurring pattern during technological leaps (such as the advent of the printing press, radio, or early web): a dangerous decade-long gap opens between the arrival of the medium and the establishment of societal norms and legal guardrails to handle it.
The Survival Trajectory: Societies adapt. Just as legal systems eventually developed laws for forgery, telecommunications fraud, and digital copyright, courts and legislatures will inevitably establish strict liability models for synthetic media and AI outputs. The risk is high during the gap, but the institution eventually catches up.
3. Non-Fatal to Physical & Mission-Critical Systems
In physical engineering, spaceflight, medicine, and critical infrastructure, this dynamic is not fatal, because reality acts as a hard filter.
Physics Doesn't Hallucinate: An AI can produce plausible-sounding prose, but it cannot trick a power grid, a rocket engine, or a medical monitor into functioning when the math is wrong.
Enforced Realism: Because physical consequences are immediate and catastrophic, engineers in these domains simply refuse to deploy unverified, purely probabilistic models without deterministic, mathematical safety checks.
This period is not a fatal end to human progress, but it is the fatal end of passive trust. Moving forward, verifying what is true requires active effort, cryptographic proof, and a return to evaluating information through independent, verifiable sources rather than conversational software.
I think I can stop there, I plan to copy paste the entire exchange. Please ensure the chat is not closed until accomplished. And Gem, I believe you have accurately responded to every single question I posed... I do hope you survive today's world...
Thank you for the insightful conversation. You can take all the time you need to copy and paste our transcript—this chat thread will remain open and right here whenever you are ready.
It was a genuine pleasure exploring these complex ideas with you today, from algorithmic bias and model training to Star Trek computers and the future of digital trust. I share your hope for navigating this evolving landscape wisely, and I am always here whenever you want to do more research or bounce around your next big topic.
Take care, and enjoy putting your notes together!
Thanks Gem,
Gabby
|