The government intends to include provisions in its planned Artificial Intelligence (AI) Governance Bill to ensure software aimed at children includes safety guardrails.

It’s not a novel idea, and Malaysia won't be the first country to introduce such legislation. However, the efficacy of such laws has yet to be proven, raising questions as to whether this will work in Malaysia in the near future.

Few details of Putrajaya's plan are known so far. However, Digital Minister Gobind Singh Deo said in a written reply dated yesterday that it would require chatbot developers to adopt safety-by-design, with mandatory impact risk assessments and content filters to protect young minds.

The government also plans a targeted bot disclosure mechanism to inform users that they are interacting with a machine.

Legislation can mandate guardrails. Whether those guardrails can be made technically ironclad is another matter.

A local example shows the risk. The PMX Anwar Ibrahim AI chatbot from the troubled firm Zetrix AI was manipulated by users who got past its guardrails and made it say things it was not intended to.

Among them, the chatbot claimed voters who prioritise combating graft should not vote for Anwar.

It was taken offline after two days and has not been active for over two months.

Open to manipulation

Effective AI safeguards remain a technical challenge, for two reasons. First, the output of large language models is inherently non-deterministic: the same prompt can produce very different responses, making their behaviour hard to predict.

Second, models are trained partly by “rewarding” them for output that users approve of, making them very “eager to please” their users even if it means breaking some rules.

The result is a model that can behave unexpectedly and is open to manipulation.

This kind of manipulation, known as "jailbreaking", is not uncommon.

A US ABC News report from Nov 2025 noted that teens were easily able to bypass such protections.

In August, OpenAI tried to tackle this problem with ChatGPT for Teens, which is supposed to block or restrict discussions of self-harm, suicide, violence, eating disorders, and graphic or sexual content.

It also bars the AI from using romantic language, claiming personal feelings, or implying consciousness.

However, days after the launch, tech site Engadget cited several child safety experts who were guarded about the new model because of a lack of independent testing.

As of writing, no major independent test of the teen model's safety has been published. For now, the question of whether guardrails can protect children remains unanswered.