A comparative experiment on youth unemployment, public policy, and the limits of AI assisted research
By Sadhana Manthapuri, PhD
We are increasingly using artificial intelligence to search for information, review literature, analyse data, and even think through policy problems.
But there is a question I have been thinking about as a researcher:
What happens when we ask different AI models the exact same policy question?
Do they give us the same answer?
And more importantly, do they tell us the same story?
I decided to test this using a policy question that is both important and complex:
How has youth unemployment in India changed between 2015 and 2025, and what policies have been introduced to address it?
I asked three AI models to explore the question: ChatGPT, Claude, and Perplexity.

The results were interesting.
The unemployment numbers were broadly similar.
The narratives were not.
Why youth unemployment?
Youth unemployment is not simply another economic indicator.
It tells us something about the transition between education and work, the ability of an economy to create opportunities for its younger population, and whether economic growth is translating into meaningful employment.
India’s youth employment question is particularly important because of the country’s large young population.
So I wanted to understand two things.
First, what happened to youth unemployment between 2015 and 2025?
Second, how did government policy respond to the problem?
I also wanted to see whether AI could distinguish between a declining unemployment rate and an actual improvement in employment outcomes.
That distinction turned out to be important.
What does the data tell us?
The first challenge appeared almost immediately.
There is no single perfectly comparable official youth unemployment series covering the entire 2015 to 2025 period.
India’s Periodic Labour Force Survey begins in 2017 to 2018, while earlier Labour Bureau estimates use a different methodology. The available datasets also differ in age groups and measurement approaches.
This means that simply putting one number from 2015 next to one number from 2025 and declaring a trend can be misleading.
Using the official PLFS series for youth aged 15 to 29, unemployment declined from 17.8% in 2017 to 2018 to around 9.9% in 2025.
At first glance, this looks like a substantial improvement.
But the broader story is more complicated.
The analysis in my original dataset showed that the decline was much stronger in the earlier part of the recovery period and became considerably slower after 2021 to 2022.
There is another important issue.
A declining unemployment rate does not necessarily mean that more young people have found good jobs.
People can leave the labour force.
They can remain in education.
They can move into informal work.
They can experience intermittent employment.
This is why I believe youth unemployment should be examined alongside labour force participation, earnings, job duration, formal employment, and the transition from education to work.
The number alone is not the story.
Then I looked at the policies
Over the last decade, India introduced a wide range of programs connected to skills, employment, entrepreneurship, apprenticeships, and formal job creation.
Some of the major interventions included:
Skill India and PMKVY
Introduced in 2015, these programs focused heavily on training, certification, reskilling, and employability.
National Apprenticeship Promotion Scheme
Introduced in 2016, NAPS attempted to strengthen the transition between education and employment through apprenticeships and employer incentives.
Startup India
Also launched in 2016, it focused on entrepreneurship and creating an ecosystem for new businesses and innovation.
PM Employment Generation Programme
Expanded opportunities for self employment and support for micro and small enterprises.
Production Linked Incentive Scheme
Introduced in 2020, it represented a broader attempt to strengthen manufacturing and encourage investment and employment.
Employment Linked Incentives
The policy direction shifted more clearly toward employer side incentives in 2024 and 2025, with the objective of encouraging formal hiring and workforce participation.
The interesting part is not simply the number of policies.
It is the direction of those policies.
For much of the decade, the policy response was strongly focused on the supply side.
Train young people.
Certify young people.
Improve skills.
Increase employability.
But what happens if the economy does not create enough appropriate jobs for those newly trained workers?
That is where the question becomes much more interesting.
And this is where the three AI models differed
I gave the same broad question to ChatGPT, Claude, and Perplexity.
ChatGPT
ChatGPT produced a relatively structured overview.
It identified the broad decline in youth unemployment, highlighted major government programs, and relied substantially on established datasets and secondary sources.
The answer was useful.
It was also easy to understand.
But I felt that it remained relatively close to the surface of the question.
It told me what happened.
It was less convincing when it came to asking why it happened and whether the policy interventions could actually be credited for the change.
Claude
Claude approached the problem differently.
The response was much more quantitative.
It examined the rate of decline over different periods and tried to connect the timing of policy interventions with observed employment outcomes.
One of the most interesting observations was that much of the decline in youth unemployment occurred before the most recent employment linked incentive programs were introduced.
That creates an important question.
Can we really attribute the decline in unemployment to those policies?
The answer is not straightforward.
Claude also raised the issue of labour force participation.
If fewer young people are participating in the labour market, unemployment can decline without a corresponding increase in employment.
That is a distinction I would want to examine carefully in any serious policy evaluation.
Perplexity
Perplexity took a broader approach.
It identified additional policies and introduced comparisons such as urban and rural youth unemployment.
Some of these additions were genuinely useful because they expanded the scope of the question.
But this is also where I became more cautious.
A number of the sources were webpages, articles, and other secondary material rather than the primary government datasets I would normally prioritise for quantitative policy research.
This does not mean the information is necessarily wrong.
It means that the source matters.
Especially when we are using the results to make policy recommendations.
So which AI model was right?
This is where my experiment became more interesting.
I don’t think there is a simple winner.
Each model gave me something useful.
ChatGPT gave me accessibility and structure.
Claude gave me stronger quantitative reasoning.
Perplexity gave me breadth and additional connections.
But each also had limitations.
And more importantly, all three models could produce a convincing answer without necessarily giving me the complete picture.
This is the part that concerns me as a researcher.
A well written answer can create an illusion of certainty.
The language can sound confident even when the underlying data have limitations.
The sources can look impressive without necessarily being the most appropriate sources.
And a model can introduce an interesting comparison that was never actually part of the original research question.
That is why I would hesitate before using an AI generated answer directly in policy research.
The data stayed the same. The story changed.
This was probably the most interesting finding from the exercise.
The three models were not necessarily disagreeing about the basic unemployment numbers.
They were interpreting them differently.
One focused on the headline trend.
Another focused on labour market participation and policy timing.
Another expanded the analysis into additional dimensions.
This made me think about something broader.
AI does not simply retrieve information. It also frames information.
And framing matters enormously in public policy.
If I tell you that youth unemployment declined from one period to another, you may conclude that the labour market is improving.
If I tell you that labour force participation has also changed, you may interpret the same decline differently.
If I then tell you that unemployment remains disproportionately concentrated among educated youth, you may ask an entirely different question.
Is the problem really a lack of skills?
Or is it a lack of suitable jobs?
That is a very different policy problem.
The question I am left with
After going through the three responses, I was less interested in asking:
Which AI model is the smartest?
I became more interested in asking:
Can AI understand the difference between a statistical trend and a policy explanation?
That distinction matters.
A correlation does not establish that a policy caused an outcome.
A declining unemployment rate does not automatically mean that employment quality improved.
A training program producing certificates does not necessarily mean that participants found sustainable employment.
And a policy being introduced before an outcome does not mean that the policy produced the outcome.
These are basic principles of policy analysis.
Yet they can easily disappear when we focus too much on getting a quick answer.
What does this mean for researchers?
I am not arguing that researchers should stop using AI.
Quite the opposite.
I think AI will become an increasingly important part of research.
It can help us identify patterns, organise literature, generate questions, compare perspectives, write code, and explore large amounts of information much faster than before.
But I think we need to change the way we use it.
Instead of asking AI:
“Give me the answer.”
Perhaps we should increasingly ask:
“What am I missing?”
“Which assumptions are you making?”
“What are the limitations of this dataset?”
“Which sources are primary sources?”
“What alternative explanations exist?”
“What evidence would I need before making this policy recommendation?”
Those questions force us to move from AI generated answers toward AI assisted research.
And there is a significant difference between the two.
My takeaway
After this small experiment, my trust in AI did not decrease.
But my understanding of where I should trust it changed.
I see AI as a very powerful research assistant.
I would not yet treat it as the researcher making the final judgment.
The models can find the numbers.
They can identify policies.
They can construct narratives.
But the responsibility to question those narratives still sits with us.
Perhaps this is the most important skill researchers will need in the age of AI.
Not simply knowing how to use AI.
But knowing when to question it.
Because sometimes the most important part of research is not the answer we receive.
It is the question we ask after receiving it.
What do you think?
Have you asked different AI models the same research or policy question?
Did they give you similar answers?
Or did you also find that the same evidence can produce very different narratives depending on the model?
I would love to hear your experience.
For the full comparative analysis, including the quantitative analysis, source evaluation, policy comparison, and my detailed assessment of ChatGPT, Claude, and Perplexity, DM me and I would be happy to share the complete analysis.
Sources and research note
The underlying analysis used the sources and datasets compiled in my original research document, including the World Bank youth unemployment indicator, India’s Periodic Labour Force Survey, ILO and Institute for Human Development reports, government policy documents, and related policy evaluations.
One important methodological caution is that the PLFS youth unemployment series is consistently available from 2017 to 2018 onward, while earlier estimates use different sources and methodologies. The figures should therefore not be treated as one perfectly comparable uninterrupted series.
The original comparison also found that the three models differed substantially in how they interpreted and framed the evidence, despite reaching broadly similar conclusions about the direction of youth unemployment trends.
Leave a comment