A community-abliterated, open-weights variant of Z.ai's GLM-4.7-Flash (30B-A3B mixture-of-experts), decensored with Heretic v1.1.0 to reduce refusals while preserving the base model's thinking mode. Text in and out, with a 200K-token context window on Venice.
A community-abliterated, open-weights variant of Z.ai's GLM-4.7-Flash (30B-A3B mixture-of-experts), decensored with Heretic v1.1.0 to reduce refusals while preserving the base model's thinking mode. Text in and out, with a 200K-token context window on Venice.
GLM 4.7 Flash Heretic is an AI model from Z.ai, available on ZeroTwo. It costs $0.07 per million input tokens and $0.40 per million output tokens, has a 200,000-token context window, can return up to 24,000 output tokens.
Intelligence
Good
Speed
Very Fast
Price
$0.07 • $0.40
Input • Output
Input
Text
Output
Text
A community-abliterated, open-weights variant of Z.ai's GLM-4.7-Flash (30B-A3B mixture-of-experts), decensored with Heretic v1.1.0 to reduce refusals while preserving the base model's thinking mode. Text in and out, with a 200K-token context window on Venice.
200,000 context window
24,000 max output tokens
Reasoning token support
Pricing
Pricing is based on the number of tokens used. For tool-specific models, like search and computer use, there's a fee per tool call. See details in the pricing page.
Priced on Venice's hosted deployment of this community fine-tune: $0.07 input, $0.04 cached input and $0.40 output per million tokens. Z.ai publishes no list price for this variant, so these rates and the 200K context and 24K output limits are Venice's, not the model author's.
Text tokens
Per 1M tokens
∙
Batch API price
Input
$0.07
Cached input
$0.04
Output
$0.40
Quick comparison
Input
Cached input
Output
GLM 4.7 Flash Heretic
$0.07
Modalities
Text
Input and output
Image
Not supported
Audio
Not supported
Endpoints
Chat Completions
v1/chat/completions
Responses
v1/responses
Batch
v1/batch
Realtime
v1/realtime
Assistants
v1/assistants
Fine-tuning
v1/fine-tuning
Embeddings
v1/embeddings
Image Generation
v1/images/generations
Image Edit
v1/images/edits
Speech Generation
v1/audio/speech
Transcription
v1/audio/transcriptions
Translation
v1/translations
Moderation
v1/moderations
Completions (legacy)
v1/completions
Features
Streaming
Supported
Function calling
Supported
Structured outputs
Not supported
Fine-tuning
Not supported
Distillation
Not supported
Fast response
Not supported
Cost efficient
Not supported
Tools
Tools supported by this model when using the Responses API.
Web search
Supported
Code interpreter
Not supported
File search
Not supported
Image generation
Not supported
MCP
Not supported
Snapshots
Snapshots let you lock in a specific version of the model so that performance and behavior remain consistent. Below is a list of all available snapshots and aliases for GLM 4.7 Flash Heretic.
olafangensan-glm-4.7-flash-heretic
GLM 4.7 Flash Heretic
Questions
How much does GLM 4.7 Flash Heretic cost?
GLM 4.7 Flash Heretic is priced at $0.07 per million input tokens and $0.40 per million output tokens, with cached input at $0.04 per million. On ZeroTwo it is included in your plan's credits rather than billed per token.
What is GLM 4.7 Flash Heretic's context window?
GLM 4.7 Flash Heretic accepts up to 200,000 tokens of context and can return up to 24,000 output tokens.
Can I use GLM 4.7 Flash Heretic for free?
You can try GLM 4.7 Flash Heretic on ZeroTwo's free plan, which includes a monthly credit allowance across every model in the catalogue. Heavier use moves to a paid plan rather than a per-token bill.