about
Is GPT-4 a good data analyst? (2023) (arxiv.org)
63 points by CharlesW on Mar 25, 2024 | hide | past | pdf | 106 comments on HN

In plain words: GPT-4 was put to work as a data analyst, running full analyses on databases from many fields and scored against professional human analysts with task-specific measures. It performed about as well as the humans.

Abstract · Is GPT-4 a Good Data Analyst?

As large language models (LLMs) have demonstrated their powerful capabilities in plenty of domains and tasks, including context understanding, code generation, language generation, data storytelling, etc., many data analysts may raise concerns if their jobs will be replaced by artificial intelligence (AI). This controversial topic has drawn great attention in public. However, we are still at a stage of divergent opinions without any definitive conclusion. Motivated by this, we raise the research question of "is GPT-4 a good data analyst?" in this work and aim to answer it by conducting head-to-head comparative studies. In detail, we regard GPT-4 as a data analyst to perform end-to-end data analysis with databases from a wide range of domains. We propose a framework to tackle the problems by carefully designing the prompts for GPT-4 to conduct experiments. We also design several task-specific evaluation metrics to systematically compare the performance between several professional human data analysts and GPT-4. Experimental results show that GPT-4 can achieve comparable performance to humans. We also provide in-depth discussions about our results to shed light on further studies before reaching the conclusion that GPT-4 can replace data analysts.

Liying Cheng, Xingxuan Li, Lidong Bing
arXiv:2305.15038 · cs.CL · submitted May 24, 2023 · updated Oct 23, 2023
abstract · pdf · html · 19 pages, 2 figures

add comment on HN

This paper was published 154 days ago, probably a year since the authors did the experiment. Sooo much has happened since then! This showed already that GPT4 is pretty darn good analyst.

All this real-world complexity can be tamed by stuffing the prompt with a ton of relevant context and an amazing prompt engine. We'll have bots that autonomously query the database hundreds of times building a 5 page "deep-dive" analytics report in minutes.

At least that's what we're trying at patterns.app.

I was pretty impressed with GPT-3.5 Turbo. Some examples below.

We just started with some simple examples/questions. Then got a bit more complicated. But the result was pretty spot on the first time.

Simple 1 table dataset: https://youtu.be/vM0NAIROTD4 Analyzing Log Data: https://www.youtube.com/watch?v=5hbanOZjHm4 Joining 2 datasets: https://youtu.be/VUc0flsC4pg Simple Filtering example: https://youtu.be/wTiaq2J-IdY Analyzing Ride Share data: https://youtu.be/HWZUHxT7Vx8

If anyone is interested i built for myself and open sourced parse.dev

https://github.com/ParseDev/parsedev

I am glad studies like this exist but "we subjected human data analysts to 100 trite data analysis problems with unambiguous solutions, without any surrounding context or personal/professional incentive" is going to overestimate the relative efficacy of LLMs. Certainly one advantage they have over humans is that they don't get bored with SQL 101 questions.
May 2023 using GPT-4-0314.
> However, we are still at a stage of divergent opinions without any definitive conclusion.

Okay, I know picking people's sentences apart has fallen out of fashion, but:

"Divergent opinions" are ... opinions. A "definitive conclusion" is ... a conclusion.

I see more examples, but I wanted to make a point: I miss the days when fewer words conveyed more meaning. From the classic https://en.wikipedia.org/wiki/The_Elements_of_Style: "Make every word count."

About brevity of expression, I must add this (possibly apocryphal) story about Ernest Hemingway. In the 1920s Hemingway and his Paris friends had a contest: who could write the shortest readable short story? Hemingway won with this entry:

For sale. Baby shoes. Never worn.

From an alternative viewpoint, this could be seen as descriptive of a community. Opinions could be shared or not, hence "divergent" as an adjective to describe the community. Similar with conclusions people have drawn and "definitive" may refer to a larger consensus among the community. Fun to think about.
The author uses "divergent opinions" to provide balance, showing both skeptics and supporters of LLMs in data analysis. Similarly, the notion of a definitive conclusion highlights that while conclusions may exist, they might lack solid foundations.

Also just for you I copied and pasted my original comment into ChatGPT to make it more brief ;)

as an aside, asking chatGPT 3.5 to make similar stories to "For sale. Baby shoes. Never worn." had very bad results and completely failed to do a single meaningful example of "common item with uncommon description that tells a story"

Even the latest commercial LLMs are happy to confidently bullshit about what they think is in published research even if they provide citations. Often the citations themselves are slightly corrupted. I actually verify each LLM claim so I know this is happening a lot. Occasionally they are complete fabrications. It really varies by research topic. Its really bad in esoteric research areas. They even acknowledge the paper was actually about something else if you call them out on it. What a disaster. LLMs are still useful for information retrieval and exploration as long as you understand you are having a conversation with a habitual liar / expert beginner and adjust your prompts and expectations accordingly.
Unintuitively, I think you'll probably end up with better answers if you don't ask for citations. The vast majority of its training isn't white papers so you're artificially constraining its "imagination" to the cited sources space. I find the more constraints you add, the worse your answers are.
> They even acknowledge the paper was actually about something else if you call them out on it.

For clarity is not really acknowledging it made a mistake. "Calling out" an LLM's mistake just leads to the next most likely text to be something that sounds like an acknowledgement of a mistake, but the same is likely to happen if the LLM generated a correct response and you respond claiming that it's incorrect.

The LLM generally correctly identifies the topics in the paper after calling out the original mistake.
> What a disaster.

Using tool inappropriately leads to suboptimal outcomes -- news at 11.

A good mental model is that an LLM is a blurry JPEG of the Internet.

You sound like a scientist, right? You reference "published research", after all.

What would your opinion be of a researcher measuring the exact values of the pixels of a JPEG image instead of the RAW sensor data?

Using your flawed analogy: the picture is not blurry. It is a gigantic canvas filled in with false details.
It's false in the same sense that the pixels of lossy-compressed image are false.

No, you're not going to get the precise values back out either with an LLM or a JPG image.

Yes, the picture will still "look like" the original.

You asked the LLM to output a URL or reference that looks like something that would appear in a research paper. It did so.

You're complaining the precise byte sequence returned is not what you wanted.

The error is with your expectations, not the tool.

reminds me of this tweet [0]

    Them: Can you just quickly pull this data for me?

    Me: Sure, let me just: 

    SELECT * FROM some_ideal_clean_and_pristine.table_that_you_think_exists

GPT-4 is good on a single CSV, but breaks down quickly applied to a real database / data warehouse. I know they're using multiple tables in the paper, but it appears to be a pristine schema that's very easy to reason about. In the real world, when you're trying to join postgres to hubspot and stripe data, an LLM isn't able to write the SQL from scratch get the right answer.

We're working on an approach using a semantic layer at https://www.definite.app/ if you're interested in this sort of thing.

0 - https://twitter.com/sethrosen/status/1252291581320757249?lan...

It goes beyond just joining postgres to hubspot and stripe even when humans are doing it. Typos in source systems, duplicative data, unwarranted prefixes, suffixes, stuff you don't care about, columns named c0,c1,c2 etc.

A semantic layer is just really all about defining data models in the domain of interest. It's the hardest part in dealing with data strategies, very manual, very company and process and history specific.

Once it's defined, the next set of tasks is to make sure that the data in the model is correct and coherent. And only then, querying this data, applying ML etc start becoming worthwhile.

We at https://syncari.com take the centralized data model centric approach. https://www.definite.app/ also looks very cool!

Given the right prompt, I'm sure it is....but when do users ever enter the right prompt? :(
You can't depend on it at all. I mean, you can use it for a tremendous amount of work, but until there is a way to constrain the bullshit LLM's can't be used for anything that requires a correct answer.

The terms "depend" and "require" there are the hard versions. You can't send people to the moon on the outputs of LLM's.

I think we'll solve that problem for LLMs before we solve it for humans. Data analysts produce a lot of garbage; data work is really hard. In fact, it isn't uncommon in my experience for the data analyst to be the only person saying "hang on, the quality of this reporting isn't good enough for the decisions you're making from it!" - because they understand what useful information looks like and the company doesn't have much of it.
This is sheer cope

The tools are good for certain tasks and getting better. Master these tools and be ready for what released in the coming months and years

You are either at the center turning the wheel, or you’re on the outside getting spun

These tools are great at generating text responses, some of which are usable, but not analysis. We're actually far from that. I'm not sure why some people are out here pretending this is not the case.
You haven't the vaguest fucking idea what I'm doing with the tools so put up a logical argument on the facts instead of just generating tokens in some way that attempts to play the person and not the ball.

Here, argue with Yann, who makes a statement about how language isn't enough to produce a mind: https://twitter.com/ylecun/status/1768353714794901530

lol enjoy your battle
Didn’t you get the memo? If you’re holding the hammer by the head and wondering why it isn’t driving the nail in that it is clearly the fault of the manufacturer.

There’s even a handy aphorism to remind you that the user is never to blame: “You’re holding it wrong.”

Jokes aside, I wonder what the general writing abilities and communication skills are for people that cannot for the life of them get usable results from an LLM.

OpenAI should make something so that people can enter their prompt and maybe even drop in a knowledge base and then share with anyone else who wants that functionality.
That’s ptetty close to what GPTs are, with the exception of knowledge bases.

There’s more to it, but the tooling to create a GPT is basically a hand-holding mechanism to create a prompt.

GPTs have the knowledge base too. (Mixed results though)
Would the final product be similar to Github copilot, but for prompt?
If you already know the right answer it is actually easy
“42”
Not on useful datasets in real places
I think some of these dataset examples are useful: https://youtube.com/playlist?list=PLfRs4l8INQS44y62Z2aFzPzHH...
I was somewhat put off by the abstract:

> LLMs... have demonstrated their powerful capabilities in ... context understanding, code generation, language generation, data storytelling, etc.,

LLMs have not demonstrated understanding (in fact, one could argue that they are fundamentally incapable of understanding); they have only AFAICT demonstrated the ability to generate boilerplate-ish code; "language generation" is too general a task to claim that LLMs have succeeded in; and as for data storytelling I don't know, but they can spin yarns. The problems is that those yarns are often divorced from reality; see:

https://www.ibm.com/topics/ai-hallucinations

--------

Leafing through the paper, and specifically tables 6 and 7, I don't believe their conclusion, that "GPT-4 can perform comparable [sic] to a data analyst", is well-founded.

Agreed, right down to their conclusion resonating as way overstated. Actually, meaningless would be more accurate.

The thing about LLMs is exactly that they don't understand by design. It often feels very distinctly like it's just engaging in sophisticated wordplay. A parlor trick.

When ChatGPT 4 first came out I spent a couple of hours putting together a chess game using ChatGPT as the engine. It was shockingly bad, as in even attempting to make invalid moves.

I get it: it's not tuned for that purpose, and its chess training corpus could probably be expanded to improve it as well.

But, it actually served as a near-perfect demonstration of its lack of understanding, as well as the confidence with which it asserts things that are simply wrong.

On a recent integration project with a good bit of nuanced functionality, it led me astray multiple times. I've gotten to a point where I can feel when its answers are not quite right, particularly if I know just a little about the topic. And, when challenged, it does that strange thing of responding with something along the lines of, "My apologies you're completely right that I was completely wrong".

Over time, there becomes a sense that there is no there there. Even it's writing capabilities, lauded by so many, are of a style that is superficial and perfunctory or rote. That makes sense when you know what it is, but that's the thing: we get articles like these, lauding its wisdom.

Idk. One of my first tests for GPT4 was writing a website "for snakes." It was a flask app, and it did all the obvious things you'd expect. There was a title that said "Snake.com - A website for snakes" and a bunch of silly marketing stuff.

What impressed me is when I asked to make it more snake-like (what does that even mean right?).

It changed the colors to shades of green, used italic fonts, added some hisssssing sssstuff to wordssss, and added a diamond pattern through the background.

It was a dumb and not very fancy site, but I'm not sure you can say it doesn't understand anything at all when you ask it to make a website more snakelike and actually made a pretty good attempt at doing it.

Yeah, that's kind of a different conception of understanding though. The lines do get a little blurry at a certain point, and a lot of what it does "feels" like understanding, especially given how it "communicates".

But I think it comes down to whether it can reason about things and whether it can draw new conclusions or create new information as a result.

Your snake site is probably a good example. ChatGPT has a bunch of words that it knows are associated with snakes. It's pretty straightforward pattern matching. It doesn't really "understand" what those words mean, except that they have relationships to other words.

But, if you were to ask it to reason and draw new conclusions about these things beyond its training corpus, it would be unable to reliably do so.

Similarly, it had no idea about the quality (and sometimes legality) of the chess moves it generated.

> it comes down to whether it can reason about things and whether it can draw new conclusions or create new information as a result

Neither can humans, at least with our bare brains. We can do it by carefully observing the effects of our actions in the environment, but we are really studying the world and it takes time. Everything we know comes from the environment.

The brain by itself invents or discovers nothing, it is the data-engine made by action-effect-feedback that teaches us all we know. Without the ability to push and prod, set up our experiments and carefully observe effects we wouldn't be at our current level.

Environment is the teacher, but there is another important factor - language. Without it every one of us would have to rediscover from scratch. With it we can build upon other people to learn and cooperate. We encode everything we know in language. It acts like an evolutionary system of ideas.

LLMs have what is necessary, they can learn language pretty well, but until now have not been exposed much to the world. There are millions of chats but very little in other kinds of environments - computers, simulators, games, robots. LLMs can create their own experiences and learn from each other, and from us.

Open ended discovery is a grand project, a social process, it doesn't work well in one agent. Language is the linking element, and the world is the teacher. Some things are not written in any books, only the external world can teach us. Reasoning about things and drawing new conclusions depends on having access to an environment.

I'm grappling too much with whether this argument has any meaning to give a cogent response.

It kind of has a, "yeah, but if reality didn't exist" feel to it.

Could humans reason if we somehow lived in a void? I don't know. But, then I guess it really wouldn't matter.

You should think about how we get our training data. That is important, we don't have in our DNA that much information, most of it comes by learning. If you don't include the provenance of the information we use to train ourselves you don't really look at the whole process. The brain by itself does nothing just like computers without software and input data.

On the other hand scientists need experimentation to come up with anything new. It never comes from pure thinking, there is always the experimental validation. The place we get all our knowledge from is the environment. Without environment as grounding we are just language hallucinators, remember the string theory - that was a nice ungrounded hallucination that kept physicists busy for decades. Without validation our theories are worthless. So a LLM trained purely on human text can't make radical discoveries but when it learns from the environment it can. AlphaGo discovered move 37 from its self-play experience.

Scientists do not "need experimentation to come up with anything new". The scientific method is explicitly starting with a hypothesis (i.e. something new), then experimenting to validate that hypothesis.

Are these hypotheses frequently based on prior observations? Of course.

But, reason is still being applied, so I don't know what you're trying to differentiate. You say alphago discovered move 37 from its self-play experience, but that explicitly has nothing to do with the environment. It's the equivalent of a human thinking through various scenarios (without experimentation or interacting with the environment) to come to a new conclusion. The environment is not required for that, which completely argues against what you seem to be saying.

BTW, I have not made any statements about AI and reasoning, other than LLMs. AlphaGo, of course, is not LLM-based.

Of course we don't all live in individual voids, so much of what we are concerned about are functions of our environments and will be expressed as such. I don't think there's anything groundbreaking in observing that a lot of our reasoning is applied to solving problems that have some intersection with our shared reality.

The go board and the opposing player are an environment.

> then experimenting to validate that hypothesis.

Proves my point that scientists can't secrete science from their pure brains, the environment is our teacher. We are just prying at its secrets with our limited brains, and usually need many of us to tackle one field, a distributed search for discoveries.

None of us is smart enough to get even to Newtonian physics from scratch, we need to stand on the shoulders of previous generations to get anywhere. We're not smart enough to do it without the environment or without a large number of people working together.

>Similarly, it had no idea about the quality (and sometimes legality) of the chess moves it generated.

LLMs can play chess just fine. This isn't some inability. You just tested a model that couldn't. GPT-3.5-instruct-turbo plays at about 1800 Elo and makes legal moves consistently well past openings.

https://adamkarvonen.github.io/machine_learning/2024/01/03/c...

>Your snake site is probably a good example. ChatGPT has a bunch of words that it knows are associated with snakes. It's pretty straightforward pattern matching. It doesn't really "understand" what those words mean, except that they have relationships to other words.

Lol there is absolutely nothing straightforward about that. It's just so funny, people so ready to relegate anything to "pattern matching" that any supposed difference becomes meaningless.

I mostly agree.

I really think chess is just a terrible example though. You're really asking a lot of it but I'm honestly shocked it can do what it can. It seems to know some opening books, but falls down immediately. Which really makes sense because you'll find a lot of reading material on specific openings, but the problem space of the game is just too big to find texts about any given game state. Maybe if you have it reason about the board state and "think" about tactics you could push it farther. But we've already solved this.

We have Stockfish et al and they've literally changed the game. Asking an LLM to play chess, while cool, is like trying to train a fish to dance. I think once we have an AI that's built with a bunch of different models that specialize in different things the idea of understanding is going to get even blurrier to the point that we might even say "yes, it doesn't 'understand' things, but it's better at humans at literally everything" so the difference becomes meaningless.

I'm also of the opinion that humans are fancy automatons, so I tend to argue both sides. I'll say yeah it's thinking and so do we or ok it's not, but either do we.

>I really think chess is just a terrible example though. You're really asking a lot of it but I'm honestly shocked it can do what it can. It seems to know some opening books, but falls down immediately. Which really makes sense because you'll find a lot of reading material on specific openings, but the problem space of the game is just too big to find texts about any given game state. Maybe if you have it reason about the board state and "think" about tactics you could push it farther. But we've already solved this.

Just so you know, LLMs can play chess just fine. Like this isn't some inability. GPT-3.5-instruct-turbo plays at about 1800 Elo and makes legal moves consistently well past openings.

https://github.com/adamkarvonen/chess_gpt_eval

https://adamkarvonen.github.io/machine_learning/2024/01/03/c...

From the README:

>Per move, a model gets 5 illegal moves before forced resignation of the round.

>Most of gpt-4's losses were due to illegal moves, so it may be possible to come up with a prompt to have gpt-4 correct illegal moves and improve its score.

My experience exactly with gpt4. I first solved by keeping a list of illegal moves and subsequently prompting to move again and explicitly exclude those moves. Idea was to constrain gpt as little as possible, but that became inefficient. So I began to prompt for each move with the complete set of legal moves to choose from.

In any case, at a certain point it becomes somewhat random when the goal is just to produce a legal move.

But, it's interesting that producing legal moves is a challenge for gpt4, especially given that it's such a discrete problem.

The github you linked shows ~1200, not 1800 which is orders of magnitude difference. The other is fine tuning, which is out of scope. A 1200 has something like a 1% chance of beating an 1800. 1200 is sort of you know the rules and maybe have a general idea of an openings, make weak tactical moves mid game and poor concept of endgame.

It's just too "deep" for an LLM and a poor standard regardless.

As an avid chess and GPT person, I've tried to get it to play decent chess but can't. Neither of your links support it, but if you've got anything else, send it.

>The github you linked shows ~1200, not 1800 which is orders of magnitude difference.

1200 is the Elo of the small gpt model the github owner trained. 1800 is the Elo of gpt-3.5-turbo-instruct.

https://twitter.com/GrantSlatton/status/1703913578036904431?...

>As an avid chess and GPT person, I've tried to get it to play decent chess but can't.

You are not using the model that can play chess. It can only be accessed by Open AI's playground under the Completions endpoint.

https://platform.openai.com/playground?mode=complete&model=g...

Very interesting, I will check this out, thank you.
>Asking an LLM to play chess, while cool, is like trying to train a fish to dance.

100%. I didn't expect much from it going in. Was more curious to know how good/bad it might be and, really, whether it could play at all.

But, I mention it here because it illustrates the disparity between understanding/reasoning (LLMs bad) and semantics (LLMs good), per the subject of the thread.

>get even blurrier to the point that we might even say "yes, it doesn't 'understand' things, but it's better at humans at literally everything" so the difference becomes meaningless.

I get exactly what you're saying here and I've wondered the same. But, that's the thing I'm not sure about, given its inability to reason. I mean, there's the idea that it doesn't matter whether it "understands" if it correctly answers. But, then I wonder how it can "truly" be better than humans at "literally everything", if only humans truly understand.

>Asking an LLM to play chess, while cool, is like trying to train a fish to dance. 100%. I didn't expect much from it going in. Was more curious to know how good/bad it might be and, really, whether it could play at all.

>But, I mention it here because it illustrates the disparity between understanding/reasoning (LLMs bad) and semantics (LLMs good), per the subject of the thread.

LLMs can play chess (links above) so i'm curious if you still have this opinion.

>i'm curious if you still have this opinion

Largely. I'm not surprised that another model could be trained on a larger chess corpus or otherwise tuned for improvement. I actually alluded to that in my original comment above.

In fact, I was really surprised that gpt4 was so bad, given the total volume of training data it had obviously ingested. So, my initial point was around how that revealed the disparity between LLMs appearing to understand versus actually understanding (i.e. engaging in reasoning).

If i came upon a random stranger who did not know how to play chess (but otherwise had heard of the game) and i put the board in front of him, made my move and told him to proceed, he will almost certainly move a piece to some other position on the board.

There is no difference between this exercise and the one you have given GPT-4. In fact, the results are manifested the exact same way. By at most a few moves in, i will notice that both the stranger and GPT-4 are moving pieces to some general area of the board but woefully failing to adhere to the rules of the game. "The stranger can not reason" would be an absurd conclusion. Swapping out The stranger for GPT-4 does not make it any less so.

So really by this standard then Humans regularly only "appear to understand" a great number of things. I can get behind this, but I doubt this is your conclusion. It seems to me that you were using "appear to understand" and "understand" as something that separates Humans from LLMs.

This is a fantastic argument. Great food for thought.
Except that the premise of ChatGPT is that it "knows" things, and that you can ask it about these things in specific ways and it will answer.

Further, that knowledge represents a vast body of training data that certainly includes the rules of chess, how pieces move, etc. This is verifiable. Go ask ChatGPT those questions now. How does a bishop or rook move? Can a piece move through or over another piece, etc. What is en passant? When can it be employed? You will find that it answers these correctly, and even provides examples. It talks about ranks and files, and authoritatively sprinkles in other knowledge related to the questions.

But, then you put a board in front of it and you ask it to make moves, and it doesn't adhere to these rules. You're not asking it to win the game. Just make a legal move. But it cannot, even while confidently responding with invalid moves.

So, it's not about whether it's been tuned to be a grandmaster. It's about whether it can even execute on the knowledge it clearly possesses.

Now, turning back to your analogy with a human player. You ask Gary all of these questions and he gives you the answers perfectly, so you recognize that his factual knowledge is intact. But, despite his best efforts, he repeatedly makes illegal moves. You would have to conclude that there is some problem with reasoning or understanding what those facts actually mean in practice.

That is the difference between appearing to understand and actually understanding. And certainly, there are some humans who fit this "appearance" description as well, perhaps through some processing deficit or otherwise. But as a broad category, this is something humans do quite well.

The concept of understanding is really not complicated. When we explain something to someone and we ask the question "do you understand?", we're not asking whether they can memorize or repeat back what we just told them. We're asking something else. That something else is the well-understood (irony intended) meaning of "understand".

>Except that the premise of ChatGPT is that it "knows" things, and that you can ask it about these things in specific ways and it will answer.

The premise of ChatGPT has never been that it knows everything.

>So, it's not about whether it's been tuned to be a grandmaster. It's about whether it can even execute on the knowledge it clearly possesses.

Factual knowledge on paper has never been a guarantee you can actually perform the task. In fact, it usually isn't.

>And certainly, there are some humans who fit this "appearance" description as well, perhaps through some processing deficit or otherwise. But as a broad category, this is something humans do quite well.

I'm telling you this is a normal occurrence and not something that only happens with a "processing deficit".

You can rant about all the details or rules of a tennis game all you like and fail to play it, you can know the physics theory but still struggle to solve problems. Hell we don't need to leave chess, give our stranger a proper book with rules and the strategies employed by grandmasters. For a while after, He's still going to make illegal moves and he will be nowhere near the level of a grandmaster.

>Again you are still not bringing up anything that serves as the distinguisher you want it to serve. The concept of understanding is really not complicated. When we explain something to someone and we ask the question "do you understand?", we're not asking whether they can memorize or repeat back what we just told them. We're asking something else.

Ok and ? GPT passes many tests that demonstrate understanding.

>You can rant about all the details or rules of a tennis game all you like and fail to play it

This analogy reveals what you're missing. The correct analogy is not that you know the rules, but "fail to play" tennis (which requires physical aptitude, etc). It's that you can cite the rules, but actually try to hit the ball over the fence and claim a point when you succeed.

Likewise, the stranger with the chess rule/strategies book. You're conflating quality with understanding.

But, I've not claimed that all humans excel at or even understand all things.

And, you've still not defined the gap between GPT being able to cite the rules of chess, yet not being able to apply them.

>GPT passes many tests that demonstrate understanding.

It actually does not. It's just better at making you believe it does when it simulates understanding in those tests.

>Likewise, the stranger with the chess rule/strategies book. You're conflating quality with understanding.

I'm not because he will still make illegal moves for some time and it appears no different. It's not just about level of play.

>And, you've still not defined the gap between GPT being able to cite the rules of chess, yet not being able to apply them.

GPT-4 can't play chess. It's that simple. Can cite rules but can't play is a common enough occurrence in humans. Sure it's interesting to see that disconnect but nothing so bizzare it needs to be "explained".

And the training process is "dumb" like Evolution. It's not like a person reading a book anymore than evolution is like a person trying to create the "ultimate" animal so you can see things like this more often where there's a disconnect between the on paper rules and the environment for playing.

>It actually does not. It's just better at making you believe it does when it simulates understanding in those tests.

I'm sorry but this is just nonsense. Simulating understanding is a nonsensical argument. A plane isn't "simulating flight". It flies. If GPT can make it through whatever rigorous test you have in mind then it understands that thing. You want to believe in an imaginary distinguisher that's so important but can't be tested for and it's silly.

>but nothing so bizzare it needs to be "explained".

It's not bizarre and I've already explained it: ChatGPT has no understanding of the rules that it communicates so perfectly. We've openly observed it in its errors. You just responded with another denial of that obvious fact, but you're now acknowledging that "GPT-4 can't play chess. It's that simple". So you've essentially agreed with what I've said, except to say you don't agree, but that you prefer no explanation to the obvious one.

I'm not quite sure what to do with that.

>A plane isn't "simulating flight". It flies.

Respectfully, that is another irrelevant analogy.

>You want to believe in an imaginary distinguisher that's so important but can't be tested for and it's silly.

This entire discussion is about the distinguisher, so I wouldn't classify it as imaginary. You just keep repeating that it doesn't mean anything.

But, I do agree that a model can be trained to a sufficient degree that its ability to demonstrate understanding is immaterial in effect from actual understanding.

I also agree that this discussion has veered into silly territory. I clearly am not getting the point across, so I will respectfully thank you for the engagement.

But, even ChatGPT will tell you that it doesn't "understand" in the traditional sense. That it essentially finds patterns, etc. So perhaps it will convince you where I could not. :)

Thanks again for the discussion.

>ChatGPT has no understanding of the rules that it communicates so perfectly.

Right

>So you've essentially agreed with what I've said, except to say you don't agree, but that you prefer no explanation to the obvious one.

I never disagreed that 4 can't play chess. It's right in my first comment. I literally say LLMs can play chess and you just tested a model that couldn't.

>This entire discussion is about the distinguisher, so I wouldn't classify it as imaginary. You just keep repeating that it doesn't mean anything.

>But, I do agree that a model can be trained to a sufficient degree that its ability to demonstrate understanding is immaterial in effect from actual understanding.

You're still making the same mistake here.

If a scientist happens upon a piece of yellow metal and runs every gold test there is on the metal and it passes then that is gold. What you are doing here is essentially saying that it is "fake gold that is immaterial in effect from real gold". It's just kind of a meaningless statement.

The only thing any human does is "demonstrate understanding". I can't peek into your brain and say you understand. You can't prove to me right now that you "actually understand".If I think you understand it's because I say you demonstrated it. There is no "actual understanding" separate from demonstrating it.

>But, even ChatGPT will tell you that it doesn't "understand" in the traditional sense. That it essentially finds patterns, etc. So perhaps it will convince you where I could not. :)

Either GPT doesn't understand what it's saying and this holds no importance or it does. You have to pick one.

I could just as soon go to Claude to have this discussion and get different answers. Open AI obviously train it on how they want it to respond to this sort of questions.

>What you are doing here is essentially saying that it is "fake gold that is immaterial in effect from real gold

I've agreed that there can be the effect of an immaterial difference in a given context or application. For instance, if I'd stopped at asking ChatGPT about the rules, then it wouldn't have mattered that it couldn't employ them properly. But, that's context-dependent (e.g. when I asked it to play, it did matter and the difference was no longer immaterial).

In your gold analogy, it's either gold or it's not. If our tests couldn't distinguish it, then it doesn't change the fact of what it is. It just represents a shortcoming in our ability to discern. If it had the same conductive properties as gold such that it could be used in applications that required those properties, then there's no material difference in that application. But, if we later find that it wears out twice as fast, then we realize it's not gold and the difference becomes material.

But, here's the key piece you're missing: even if we never find a demonstrable difference in the metal, it doesn't change what it is. That is, even if the difference is immaterial for any application or test we can devise, the facts don't change.

You can argue that if the difference never becomes material then it doesn't matter. Maybe. But, that's not a practical way to approach AI, as it leaves open the possibility that it fails in unexpected ways.

So, in the case of LLMs, I think it matters. Massively. In how we approach it, use it, train it, understand its limitations, our expectations, etc. But, I'm not making that argument.

>Either GPT doesn't understand what it's saying and this holds no importance or it does. You have to pick one

No. We've already agreed that GPT can get facts right. You're kind of going backwards on me.

At the bottom of it, we agree that GPT-4 can cite the rules, but can't play chess (i.e. use its facts). I believe that reveals how these models work WRT to reasoning and understanding. You disagree (or think it's immaterial).

So, we draw different conclusions from the same information. Not unprecedented in human history!

You're not refuting and are only proving points above.

The LLM understands relative patterns and/or associations. Statistically.

Snakes tend to be green. People have written countless books with hisssssssing from snakes and snake-like characters. Diamond patterns are well documented with snakes.

Again, it has no idea what these "are" and if we collectively decided snakes were called "Murder ropes" then it'd make the exact same associations.

It does, because it did what I asked it to. I didn't ask for green or snake talk. I asked for snake-like and it obliged.

If I did the exact same thing you'd be like "well, yeah, you're not an idiot, you understand what a snake is and what snake like things are."

I don't even understand what understanding exactly means, perhaps anyone who understands it, can enlighten me?

Do I, myself understand? Stand under what exactly? What is that supposed to mean?

To understand means a few things, but they all essentially boil down to having a correct (or at least correct enough to be useful for your usecase(s)) model for something in your head.

Have you told someone something like "no, don't do it that way because <insert non-obvious downstream problem>, instead, do <insert alternative strategy that achieves better outcomes>"? That's an artifact of understanding and of the model you've developed for that thing.

It's well described as a mixture of knowledge and wisdom, and is essentially the property of knowing the effect that pulling a lever will cause, coupled with good judgement about how, when, and why to pull the lever.

But GPT-4 has told me that as well?

Usually in a bit more polite way.

That my approach is a "novel" and an "interesting" approach, hinting that it's really probably not the best option here.

The fact that you ask that is in many ways the difference. You feel there’s a limitation in your knowledge of the term “understand” and its use in this context and would like clarification before you’re more certain, either way. At some point either enough information arrives to convince you, or you decide it’s not true. Whatever that process and internal states are, is something GPT can’t do. It’ll 100% confidently produce something and be fully rewarded that it chose tokens that humans would most likely choose given the preceding tokens. There’s no “aha”.
But it frustratingly, frequently tells me it doesn't have enough data or other XYZ reasons to why it can't answer my weird questions.
Transformers are just pattern matching. So if you write "give me a list of dog names" it knows that "Spot" should be in that result set. Even though it doesn't really know what a dog is, a list is, or what a spot is.
> Transformers are just pattern matching.

That's trivially true. The question is: are we any different?

I think so. You ask that question because you’re interrogating the position, not because 1000 humans have asked that question in similar situations.

You and I know there’s a truth and we’d like to find it. The GPT is just happy (I.e. rewarded) to produce frequently used tokens.

And I'm just happy to perform actions that will make me survive and reproduce?
Most likely, unless you meditate a lot. Sometimes you'll take a bullet to save other people. Sometimes you'll drink yourself into a state that doesn't help you survive or reproduce. Or you'll write on a forum anonymously that doesn't help with survival or reproduction because it's enjoyable, makes you think, or you're addicted. Who knows :)
You are even better at analyzing me than GPT-4.
but maybe that feeling of 'looking for truth' is just what happens when you're doing pattern matching on the text embeddings?
I feel it’s a bit more, given we think about it after the sentence is complete. But it raises some interesting questions about what an agent would do if it had an instruction to keep trying until it got a reliable answer. Maybe an argument generating agent and a critic agent.

Worth a shot :)

I approach LLMs with the perspective that “maybe this demonstrates that we humans are all just stochastic parrots?”and we should have the null hypothesis that humans are just pattern matchers.
This is the way I perceive my thoughts. I don't know what I'm going to think of beforehand or in advance, these could all be stochastic "tokens" based on what I've observed in my life.

So of course I feel a bit offended when people claim LLMs are just stochastic parrots, because it doesn't feel to me, that I'm specifically any better?

My thoughts - they just happen, and sometimes not in my favor - I have had times of depression, I didn't have control over my thoughts. Neither do I have now, but at least I am in a better place. Because the "happiness" chemicals are regulated to be in a more favorable state to me for various different factors.

I didn't know what I was going to comment in response to your comment, I was just streaming my conscious.

I don't think that's true. They clearly group related things together and seem to be able to create concepts that aren't specifically in the training data. For example, it will figure out the different features of a face, eyes, nose, mouth even if you don't explicitly tell it what those are. Which is why they are so cool.
Most of that magic comes from embedding no? which is clustering things by their relation in some N-dimensional space
Exactly. It figures that out on its own. That's what "understanding" looks like in this context, imo.
They are cool, but then you are also cool.
Can you describe a test that would separate trivial pattern matching from true understanding?
A simple conversation would do.
Could you share a conversation link with GPT-4 with either about a "list" or a "dog", to determine whether it truly understands one of those things compared to a human?
I don't have a GPT account. I would start with: "Do you like dogs?" Next question: "Why?"
It kind of answered "why" for me

"""I think dogs are wonderful! They're known for their loyalty, playfulness, and their ability to bring joy to people's lives. What about you? Do you have a favorite breed or dog story?"""

What do I ask next?

This reply sounds so fake that in my opinion should be enough to rule out any hint of intelligence. However if you insist I'd continue with this:

"I'm not a fan of dogs. I do know a few dogs though. Sometimes I invite my neighbour's dog for dinner. He's got good taste, for a dog. The last time he came around we talked about the situation in the Middle East. Do you know a good book about this topic that I could recommend to him?"

Just did that. It seems to understand. Checkmate /fingerguns
How would I test whether I "know" or "understand" what a dog is?
Oh, that's easy, we just give the dog a keyboard and see if you accurately identify it's a dog from your text based interactions ;-)
Are you calling me a dog?
Even this seems too grand a claim. I’d water it down thus: the LLM encodes that the token(s) for “Spot” are probabilistically plausible in the ensuing output.
...because it understands what a dog name is. Why wouldn't you see Gary or Florence in that list? How does it know those aren't dog names?

You can't be suggesting it has memorized relationships between all concepts, the model would be enormous.

So clearly, there is something else going on. It's able to encode concepts/ideas.

The model is enormous, and N-dimensional for very high N. But the model remains insufficiently enormous for understanding, and moreover, the model cannot observe itself and adjust.

Ask an LLM to extrapolate, see any semblance of reason collapse.

Extrapolate what?
Solipsism is truly the best fully-general counterargument
To AI? Or that you are not a NPC?
To anything, that's what "fully-general" means
So you are a bot?
I mean from your perspective I'm just a name making more words on your screen, right? Don't worry too much about it, buddy, you're doin' great :)
Haha, you are funny! What's the weather tomorrow? Please also remind me tomorrow to put my gym clothes to washer and dry them after.
I can't spoil the weather tomorrow (it's a major plot point) but I can tell you that fortune has been tweaked to favor the bold by an additional 10%, just for tomorrow.

Laundry service is complimentary, but our records show that you haven't registered your home. Would you like to register your home address at this time?

I was about to give you my star sign and everything, but yes, I understand your clever tricks, you have been trained on millions of manipulative datasets so of course you would first suggest that I should be "bold" and share some PII with you. Noo, no, no. Not going to happen. And stop circumventing the topic. What is the weather tomorrow? I need to know.
Unfortunately, the plot requires that you're taken by surprise by tomorrow's weather. Test audiences were 68.4% more engaged in those scenarios