സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ്
പുതിയത്Cheap multimodal reasoner with very low time-to-first-token.
റിലീസ് തീയതി: 2026 മേയ് 28
₹110.40
10 ലക്ഷം ഔട്ട്പുട്ട് Tokens-ന്
ഇൻപുട്ട്: 10 ലക്ഷം Tokens-ന് ₹19.20
0% മാർക്ക്അപ്പിൽ, ഒരു ഡോളറിന് ₹96 എന്ന നിരക്കിലാണ് ഈടാക്കുന്നത്.
നിരക്കുകൾ
ബെഞ്ച്മാർക്കുകൾ
റേറ്റിംഗ് എന്ന് രേഖപ്പെടുത്താത്ത സ്കോറുകൾ ശതമാനത്തിലാണ്. എല്ലാ ബെഞ്ച്മാർക്കുകളും സ്വതന്ത്രമായി പരിശോധിച്ചവയാണ്.
80.9%
GPQA Diamond
GPQA Diamond - graduate-level science Q&A
21.4%
HLE
Humanity's Last Exam
43.9%
SciCode
SciCode - scientific code generation
67.3%
IFBench
IFBench - precise instruction following
73.7%
Long Context
Long Context Reasoning - reasoning over long inputs
35.6%
Terminal-Bench Hard
Terminal-Bench Hard - agentic terminal tasks
39.3%
Terminal-Bench 2
Terminal-Bench 2.1 - agentic terminal tasks, second edition
സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ്-നെ കുറിച്ച്
StepFun's sparse Mixture-of-Experts vision-language model, pairing a 196B-parameter language backbone with a vision encoder for native image and video understanding while activating roughly 11B parameters per token. It is engineered for high-frequency production workloads and reaches up to 400 tokens per second. StepFun built it for production agents spanning agentic, coding, search and multimodal work, with three selectable reasoning levels to trade speed and cost against depth.
StepFun organises the release around four claims rather than one headline: multimodal understanding that turns into action, web and visual search that reaches further, tool use and orchestration that stays coherent however long a run gets, and compatibility with the agent harnesses teams already use. The last one is not filler. The lab publishes a per-harness breakdown of its own coding benchmark across six popular agent frameworks, where the average rises from 56.5% for Step 3.5 Flash to 67.1%, and, the part it actually argues for, the spread between the best and worst harness narrows sharply. A model that only performs well on one scaffold is not usable in a production stack that runs several.
The efficiency case is made in cost rather than only in speed. StepFun reports 76.5% on SWE-bench Verified, 56.3% on SWE-bench Pro (five points up on the previous Flash generation) and 59.6% on Terminal-Bench 2.1 (up 6.1). It also ships an advisor arrangement in which this model drives the whole trajectory - calling tools, reading results, iterating - and consults a larger model only at the few inflection points where its own judgement falls short, such as planning or recovery from repeated failure. With that on, the lab measures 76.3% on SWE-bench Verified at 19 cents per agentic task against Claude Opus 4.6 at 78.7% for $1.76, or roughly 97% of the score for about a ninth of the cost.
Search is treated as part of reasoning rather than as an add-on, on the explicit reasoning that a model this size cannot hold the world in its weights and should instead be good at going to get things. StepFun reports 47.2% on Humanity's Last Exam with tools against 35.7% for Step 3.5 Flash, 75.8% on BrowseComp, 92.8% F1 on DeepSearchQA and 71.7% on ResearchRubrics, which it places ahead of GPT 5.5 at 61.5% and near Claude Opus 4.7 at 73.9%.
Vision follows the same logic. Rather than growing parametric knowledge, the model is trained to reach for visual tools: with a visual search tool it reports 79.2% on SimpleVQA, matching or beating models several times its size on long-tail entity recognition, and given a code tool that lets it crop, zoom and draw on an image it reports 95.3% on the V-star fine-grained perception set and 89.1% on the 4K high-resolution variant. StepFun extended it to phone interface operation as well, at 61.9% on an Android daily-task benchmark. The behaviour it found most interesting was one nobody trained for: after writing frontend code the model opened the interface, exercised the page it had just built and revised its own code from what it saw.
StepFun's own comparison table is where the trade is visible. Its 59.6% on Terminal-Bench 2.1 sits well behind Gemini 3.5 Flash at 76.2% and GPT 5.5 at 82.7%, and on professional-deliverable scoring it is behind the frontier cohort too. That is the Flash bargain stated plainly: agent reliability, perception and search at a rate that lets you run many concurrent agents, not the best score on any single hard benchmark. The weights are published in several precisions and run under vLLM, SGLang, transformers and llama.cpp, small enough to fit a single high-memory workstation, so this is a model you could host yourself; the copy served here is StepFun's.
ലോഞ്ചിംഗ് വേളയിൽ StepFun വ്യക്തമാക്കിയത്
- Multimodal understanding and action
- It reads product interfaces, documents, charts and natural scenes, then writes code or calls tools to act on what it saw, rather than only describing the image back to you.
- Wider and deeper search
- StepFun reports 47.2% on Humanity's Last Exam with tools, 75.8% on BrowseComp and 92.8% F1 on DeepSearchQA, treating retrieval as part of reasoning rather than as an external add-on.
- Reliable tool orchestration
- The lab reports 67.1% on the ClawEval agent-reliability set, ahead of every other model in its own table, arguing for less drift, fewer broken tool calls and fewer failed runs on long tasks.
- Works across harnesses
- On the lab's own coding benchmark run through six different agent frameworks, the average rises to 67.1% from 56.5% for the previous Flash, and the gap between best and worst harness narrows sharply.
- Advisor mode economics
- With a larger model consulted only at planning and recovery points, StepFun measures 76.3% on SWE-bench Verified at 19 cents per task against Claude Opus 4.6 at 78.7% for $1.76.
- What it is not for
- StepFun's own table puts it behind Gemini 3.5 Flash and GPT 5.5 on Terminal-Bench 2.1 and behind the frontier cohort on professional deliverables. It buys throughput and reliability, not peak scores.
Frequently Asked Questions
സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ് സംബന്ധിച്ച പ്രധാന ചോദ്യങ്ങളും ഉത്തരങ്ങളും.
സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ് എപ്പോഴാണ് റിലീസ് ചെയ്തത്?
StepFun 2026 മേയ് 28-ൽ സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ് പുറത്തിറക്കി.
സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ് വികസിപ്പിച്ചത് ആരാണ്?
സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ് വികസിപ്പിച്ചത് StepFun ആണ്. 99Models AI ഒറിജിനൽ പ്രൊവൈഡർ നിരക്കിൽ തന്നെ ഇതിലേക്ക് കണക്ട് ചെയ്യുന്നു.
സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ് എത്രത്തോളം കാര്യക്ഷമമാണ്?
ഇന്റലിജൻസ് റാങ്കിംഗിൽ 54 മോഡലുകളിൽ 43-ാം സ്ഥാനത്താണ് ഇത്. ബെഞ്ച്മാർക്ക് സ്കോറുകൾ മുകളിലുള്ള പാനലിൽ കാണാം.
സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ് ഉപയോഗിക്കാൻ എത്ര ചെലവാകും?
10 ലക്ഷം ഇൻപുട്ട് Tokens-ന് ₹19.20 രൂപയും ഔട്ട്പുട്ടിന് ₹110.40 രൂപയുമാണ് അധിക മാർക്ക്അപ്പില്ലാത്ത നിരക്ക്. സബ്സ്ക്രിപ്ഷനില്ല, ഉപയോഗിക്കുന്നതിന് മാത്രം പണമടയ്ക്കുക.
സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ് API-യുടെ ഡോളർ നിരക്ക് എത്രയാണ്?
10 ലക്ഷം ഇൻപുട്ട് Tokens-ന് $0.20 ഡോളറും 10 ലക്ഷം ഔട്ട്പുട്ട് Tokens-ന് $1.15 ഡോളറുമാണ് നിരക്ക്. ഈ പേജിലെ രൂപ നിരക്കുകൾ ഡോളറിന് ₹96 എന്ന നിരക്കിൽ മാറ്റിയതാണ്.
സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ്-ന് എത്ര നീളമുള്ള സംഭാഷണം ഓർത്തുനിൽക്കാനാകും?
ഇതിന്റെ Context window 2.6 lakh Tokens ആണ്. ഒരുമിച്ച് നൽകുന്ന സംഭാഷണങ്ങളും ഫയലുകളും ഉൾപ്പെടെ ഇതിൽ വായിക്കാൻ സാധിക്കും.
സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ് മികച്ച വാല്യൂ നൽകുന്ന ഒന്നാണോ?
പെർഫോമൻസും നിരക്കും അടിസ്ഥാനമാക്കിയുള്ള വാല്യൂ റാങ്കിംഗിൽ 54 മോഡലുകളിൽ 32-ാം സ്ഥാനത്താണ് ഇത്.
മറുപടി നൽകുന്നതിന് മുൻപ് സ്റ്റെപ്പ് 3.7 ഫ്ലാഷ് Reasoning നടത്തുമോ?
ഡിഫോൾട്ടായി ഓൺ ആണ്. Reasoning പിന്തുണയ്ക്കുന്ന മോഡലുകളിൽ ചിന്തിക്കുന്നതിന്റെ വ്യാപ്തി മെസ്സേജ് ബോക്സിൽ ക്രമീകരിക്കാം.
കാറ്റലോഗ് പുതുക്കിയത്: 2026 സെപ്റ്റം 9