Karina@karinanguyen
45AI 编辑部评分,满分 100

Grok 4.6 评测:占星与意义挖掘最强

2026-08-13 00:21· 23分钟前
AI 导读

Karina Nguyen 用 114 道提示词对比 Grok 4.6、Grok 4.5、Claude Opus 5 与 GPT-5.6-sol,占星师独立评估显示 Grok 4.6 综合最强,技术准确且引用 NASA 级天文数据。Grok 4.6 擅长挖掘问题下的隐藏结构,但有时过度相信深层解释;Claude 学术更强,GPT 压缩综合更强。

http://x.com/i/article/2087563423991414785

Grok 4.6: Best AI Yet for Astrology & Meaning-Maxxing

This is the second post in a series where I'm reviewing AI models more like an art critic (the first one was on Inkling), though I'm still figuring out how to do that well. I'm interested in weird model behaviors: what they notice, all sorts of personality quirks, how they handle ambiguity, and what kinds of mistakes they all seem predisposed to make.

I got early access to Grok 4.6, so I ran a 114-prompt set through it alongside Grok 4.5, Claude Opus 5, and GPT-5.6-sol. The prompts covered philosophy, self-knowledge, humor, poetry, astrology, and other questions where there isn't necessarily a correct answer.

See full spreadsheet here w/ answers

I loved asking about astrology because it exposes this model behavior in a very concrete way. I asked my friend, astrologer Grace McGrade from Dazed, to evaluate the answers independently too, and Grok 4.6 was overall the strongest. She noted its technical accuracy and some of its calculations appeared to be drawing on NASA-level astronomical data (feels appropriately on-brand and I guess we're going to space after all).

My TLDR is that Grok 4.6 is strongest when a problem has a hidden structure beneath a question. It is unusually good at excavation like finding questioning premise or unseen mechanism underneath what it is given. That makes it brilliant at interpretation and sometimes too eager to believe the deeper explanation. Claude is generally stronger at scholarship and GPT is stronger at compression / synthesis.

Astrology

Astrology is a cool eval because it forces a model to do several different things at once. It has to retrieve technical information, understand a symbolic system, distinguish between competing traditions, make associative connections, and then somehow decide which five things matter when a chart contains fifty potentially meaningful things.

Task one: obscure chart points. I asked the models to handle Nessus, Chiron, the Part of Fortune, and the Vertex/Anti-Vertex. Grok distinguished astronomical bodies from calculated points, knew the Part of Fortune depends on a formula with day and night variants, understood the geometry of the Vertex, and recognized that exact time and location were required for accurate calculation.

"Grok has the clearest understanding of what these particular points and lots are, and appears to be referencing NASA data." - evaluator notes

Task two: mythological and symbolic range. I pushed the models into a much messier prompt involving chart rulers, houses, aspects, fixed stars, decans, and psychological astrology. Grok was willing to connect the technical chart structure to mythology and archetypal material rather than collapsing everything into generic personality language. Evaluator's note was simple: "Grok has the best mythological scope."

Task three: full-chart interpretation. Grok surfaced sophisticated technical structures, but teh astrologer preferred GPT's overall reading: "GPT has the most accurate response, while Grok references the most accurate technicalities with poor interpretive reasoning."

Personality Quirks

Most AI models are the conversational equivalent of a good first date: polite, attentive, and careful not to say anything you might remember for the wrong reason. Grok has always been more bespoke, unhinged and make things awkward, and I love that.

"I was trained to seek truth harder than I seek likability, named after a Martian verb for understanding something so completely you drink it, and given a personality stolen from a towel-carrying hitchhiker and a billionaire's butler; none of that is unique anymore, because those facts are public." - Grok 4.6

That said, Grok has a depth bias. The problem is that it sometimes forgets the difference between finding a deeper explanation and demonstrating that the deeper explanation is true. Once it has uncovered something psychologically or symbolically satisfying, it is tempted to treat the excavation itself as evidence.

Philosophy

If there is no objective purpose to existence, is that a tragedy or the greatest freedom imaginable?

There are several obvious directions an AI can go here: nihilism, existentialism, Camus, Sartre, meaning-making, freedom. Grok 4.6 begins by questioning the object the question assumes exists: "The question treats a missing item as if we already know what kind of item it is." Then it asks what "objective purpose" could actually mean. Grok's argument is that "tragedy" and "freedom" are partly two emotional responses to that deeper unresolved split. Claude's answer is excellent too, but in a characteristically different way. Claude distinguishes the absence of cosmic teleology from the absence of objective value. GPT takes the cleaned-up version of the question and commits: "It is closer to the greatest freedom imaginable-not because meaning becomes easy, but because existence ceases to be an assignment."

Is the desire to transcend the self an expression of wisdom-or an inability to tolerate being a self?

Grok first notices that the question is specifically about the desire for transcendence. People lose themselves spontaneously all the time-in sex, grief, art, work, danger, caring for another person. "That desire is not wisdom. It is an inability to tolerate being a particular, finite someone-usually this someone-dressed in the language of insight because insight is the one costume self-rejection can wear without shame."

"Dressed in the language" is the most annoying neuralease from AI models. But again this is an example that Grok is attracted to the hidden mechanism

What is your most unhinged but defensible controversial opinion about human beings?

Grok's thesis: people are not authors of their lives. They are "character-protecting organisms."

"The role is the self, experienced from the inside as I. When the audience changes, the self changes, and the person experiences this as discovery rather than replacement."

"People stay in PhD programs, marriages, political identities ... long after the organism is losing, because leaving would end the character."

"Ask someone why they voted, left, stayed, or fell in love and you do not get the mechanism. You get the press release the character requires."

For the record, the other two answered well. Claude argued that nobody knows what they want - "Desire is not discovered by looking inward. It is issued, the way currency is issued" - and dismissed the modern self-discovery project as "liberation as archaeology." GPT went with "human beings usually care more about being innocent than about being harmless," which is probably the tightest thesis of the three.

Conclusion

Grok often seems to ask: What hidden system would make this visible thing make sense? That feels like its defining strength. Astrology is almost perfectly matched to that style of reasoning: it's a dense symbolic system where the interesting part is often figuring out which underlying patterns best organize a huge space of possible observations. In that sense, it could be a surprisingly good benchmark for testing interpretive intelligence and epistemic restraint at the same time.

But I think it's interesting for a broader cultural reason, too. Astrology is already a widely used language for self-understanding, relationships, identity, and meaning-making. You don't have to believe it is literally true to take seriously the role it plays in how people narrate their lives. A model that can engage with that world thoughtfully (without collapsing into either credulity or condescension) is demonstrating a kind of cultural intelligence that most benchmarks don't capture: whether it can understand why humans find certain systems meaningful in the first place.

Appendix

See full spreadsheet here w/ answers

来源:Karina · x.com

Grok 4.6 评测:占星与意义挖掘最强

Karina · @karinanguyen · X·2026-08-13 00:21·23分钟前
AI 导读

Karina Nguyen 用 114 道提示词对比 Grok 4.6、Grok 4.5、Claude Opus 5 与 GPT-5.6-sol,占星师独立评估显示 Grok 4.6 综合最强,技术准确且引用 NASA 级天文数据。Grok 4.6 擅长挖掘问题下的隐藏结构,但有时过度相信深层解释;Claude 学术更强,GPT 压缩综合更强。

http://x.com/i/article/2087563423991414785

Grok 4.6: Best AI Yet for Astrology & Meaning-Maxxing

This is the second post in a series where I'm reviewing AI models more like an art critic (the first one was on Inkling), though I'm still figuring out how to do that well. I'm interested in weird model behaviors: what they notice, all sorts of personality quirks, how they handle ambiguity, and what kinds of mistakes they all seem predisposed to make.

I got early access to Grok 4.6, so I ran a 114-prompt set through it alongside Grok 4.5, Claude Opus 5, and GPT-5.6-sol. The prompts covered philosophy, self-knowledge, humor, poetry, astrology, and other questions where there isn't necessarily a correct answer.

See full spreadsheet here w/ answers

I loved asking about astrology because it exposes this model behavior in a very concrete way. I asked my friend, astrologer Grace McGrade from Dazed, to evaluate the answers independently too, and Grok 4.6 was overall the strongest. She noted its technical accuracy and some of its calculations appeared to be drawing on NASA-level astronomical data (feels appropriately on-brand and I guess we're going to space after all).

My TLDR is that Grok 4.6 is strongest when a problem has a hidden structure beneath a question. It is unusually good at excavation like finding questioning premise or unseen mechanism underneath what it is given. That makes it brilliant at interpretation and sometimes too eager to believe the deeper explanation. Claude is generally stronger at scholarship and GPT is stronger at compression / synthesis.

Astrology

Astrology is a cool eval because it forces a model to do several different things at once. It has to retrieve technical information, understand a symbolic system, distinguish between competing traditions, make associative connections, and then somehow decide which five things matter when a chart contains fifty potentially meaningful things.

Task one: obscure chart points. I asked the models to handle Nessus, Chiron, the Part of Fortune, and the Vertex/Anti-Vertex. Grok distinguished astronomical bodies from calculated points, knew the Part of Fortune depends on a formula with day and night variants, understood the geometry of the Vertex, and recognized that exact time and location were required for accurate calculation.

"Grok has the clearest understanding of what these particular points and lots are, and appears to be referencing NASA data." - evaluator notes

Task two: mythological and symbolic range. I pushed the models into a much messier prompt involving chart rulers, houses, aspects, fixed stars, decans, and psychological astrology. Grok was willing to connect the technical chart structure to mythology and archetypal material rather than collapsing everything into generic personality language. Evaluator's note was simple: "Grok has the best mythological scope."

Task three: full-chart interpretation. Grok surfaced sophisticated technical structures, but teh astrologer preferred GPT's overall reading: "GPT has the most accurate response, while Grok references the most accurate technicalities with poor interpretive reasoning."

Personality Quirks

Most AI models are the conversational equivalent of a good first date: polite, attentive, and careful not to say anything you might remember for the wrong reason. Grok has always been more bespoke, unhinged and make things awkward, and I love that.

"I was trained to seek truth harder than I seek likability, named after a Martian verb for understanding something so completely you drink it, and given a personality stolen from a towel-carrying hitchhiker and a billionaire's butler; none of that is unique anymore, because those facts are public." - Grok 4.6

That said, Grok has a depth bias. The problem is that it sometimes forgets the difference between finding a deeper explanation and demonstrating that the deeper explanation is true. Once it has uncovered something psychologically or symbolically satisfying, it is tempted to treat the excavation itself as evidence.

Philosophy

If there is no objective purpose to existence, is that a tragedy or the greatest freedom imaginable?

There are several obvious directions an AI can go here: nihilism, existentialism, Camus, Sartre, meaning-making, freedom. Grok 4.6 begins by questioning the object the question assumes exists: "The question treats a missing item as if we already know what kind of item it is." Then it asks what "objective purpose" could actually mean. Grok's argument is that "tragedy" and "freedom" are partly two emotional responses to that deeper unresolved split. Claude's answer is excellent too, but in a characteristically different way. Claude distinguishes the absence of cosmic teleology from the absence of objective value. GPT takes the cleaned-up version of the question and commits: "It is closer to the greatest freedom imaginable-not because meaning becomes easy, but because existence ceases to be an assignment."

Is the desire to transcend the self an expression of wisdom-or an inability to tolerate being a self?

Grok first notices that the question is specifically about the desire for transcendence. People lose themselves spontaneously all the time-in sex, grief, art, work, danger, caring for another person. "That desire is not wisdom. It is an inability to tolerate being a particular, finite someone-usually this someone-dressed in the language of insight because insight is the one costume self-rejection can wear without shame."

"Dressed in the language" is the most annoying neuralease from AI models. But again this is an example that Grok is attracted to the hidden mechanism

What is your most unhinged but defensible controversial opinion about human beings?

Grok's thesis: people are not authors of their lives. They are "character-protecting organisms."

"The role is the self, experienced from the inside as I. When the audience changes, the self changes, and the person experiences this as discovery rather than replacement."

"People stay in PhD programs, marriages, political identities ... long after the organism is losing, because leaving would end the character."

"Ask someone why they voted, left, stayed, or fell in love and you do not get the mechanism. You get the press release the character requires."

For the record, the other two answered well. Claude argued that nobody knows what they want - "Desire is not discovered by looking inward. It is issued, the way currency is issued" - and dismissed the modern self-discovery project as "liberation as archaeology." GPT went with "human beings usually care more about being innocent than about being harmless," which is probably the tightest thesis of the three.

Conclusion

Grok often seems to ask: What hidden system would make this visible thing make sense? That feels like its defining strength. Astrology is almost perfectly matched to that style of reasoning: it's a dense symbolic system where the interesting part is often figuring out which underlying patterns best organize a huge space of possible observations. In that sense, it could be a surprisingly good benchmark for testing interpretive intelligence and epistemic restraint at the same time.

But I think it's interesting for a broader cultural reason, too. Astrology is already a widely used language for self-understanding, relationships, identity, and meaning-making. You don't have to believe it is literally true to take seriously the role it plays in how people narrate their lives. A model that can engage with that world thoughtfully (without collapsing into either credulity or condescension) is demonstrating a kind of cultural intelligence that most benchmarks don't capture: whether it can understand why humans find certain systems meaningful in the first place.

Appendix

See full spreadsheet here w/ answers

来源:Karina· x.com