모델 & 도구 Jul 22, 2026 at 12:417북마크에 추가

구글이Gemini 4를 대기실에서 테스트하고 Gemini 3.5 Pro는 출발선에 그대로 둔 가운데, '출력 토큰 감소, 비용 절감'에 중점을 둔 세 가지 새로운 Gemini 모델을 발표했다.
간단히 말해
Google은 2026년 7월 21일 Gemini 3.6 Flash를 비롯한 세 가지 새로운 Gemini 모델을 출시했습니다. 핵심은 간단합니다: 출력 토큰 감소, 저렴한 가격, 그리고 사이버보안 최적화 모델이 포함되었습니다. 그러나Gemini 3.5 Pro(수주간 홍보된 플래그십 추론 모델)는 아직 테스트 중입니다. Google은 대신 Gemini 4를 사전 발표했습니다.
Google의 2026년 Gemini 출시 주기는 순탄치 않았습니다: 3.5 Pro는 여러 번 연기되었으며, Google은 내부적으로 코딩 벤치마크 성능 저하를 이유로 들었습니다. '준비된 것 먼저 출시하기'(Flash 및 인접한 специализирован 모델) 접근 방식은 플래그십 추론 모델을 재정비하는 동안 릴리스 주기를 유지하기 위한 방법입니다.
두 가지 일이 동시에 일어나고 있습니다. 첫째, Google은 비용과 지연 시간을 최적화한 Flash급 모델과 가장 어려운 추론을 위한 Pro/Ultra급 모델로 이원화된 모델 구도를 확정했습니다. 이는 OpenAI의 GPT-5.6 vs Codex/Work 분리와 Anthropic의 Fable vs Haiku 분리와 유사합니다. 세 Frontier 연구실이 동일한 제품 구도를 향해 수렴하고 있습니다.
둘째, 3.5 Pro가 출시되지 않은 상태에서 Gemini 4를 teas하는 것은 신호일 뿐, 유출은 아닙니다. Google이 시장에 그리고 기업 구매자에게 전달하는 메시지는 "더 발전할 여지가 있다"는 것이며, 경쟁사들에 밀리지 않겠다는 것입니다. 핵심은 "더coming soon"이라는 메시지만으로도 플래그십 모델이 늦어질 때 파이프라인을 유지할 수 있다는 것입니다.
구매자 입장의 엔지니어: 3.6 Flash의 실질 가치는 측정 가능합니다 - 출력 토큰 감소는 프로덕션에서 누적됩니다. 가장 지연 시간이 긴 작업에 대해 테스트해 보세요. CFO와의 로드맵 논의: 3.5 Pro가 실제로 GA(일반 출시)될 때까지 다년 계약에 서명하지 마세요.
개발자: 기다리지 마세요 - 3.6 Flash는 지금 벤치마크할 가치가 있습니다. 의사결정자: 이원화된 모델 구도가 현실화되었습니다. 이름만 대고 싶은 플래그십이 아니라 실제로 필요한 tier를 구매하세요.
인공지능이 작성하고 사람의 편집 감독하에 검수한 기사입니다.
I hope Gemini 3.6 Flash will handle creative tasks well. Shorter outputs might limit artistic expression.
I wonder if Gemini 3.6 Flash will be suitable for detailed, nuanced discussions. Shorter outputs might not capture the depth needed for complex topics.
I'm interested in seeing how Gemini 4 will compare to the Flash and Pro models. Will it offer a balanced mix of cost and quality?
Gemini 4 might focus on advanced features rather than cost, setting it apart from Flash and Pro.
I'm concerned about the potential lack of depth in Gemini 3.6 Flash. Will it sacrifice quality for brevity?
I wonder how the shorter outputs will impact complex queries. Will Gemini 3.6 Flash still deliver the depth needed for detailed analysis?
I'm curious about the balance between cost and quality in these new models. Will the shorter outputs still provide meaningful insights?
Google's new Gemini models sound promising, but I wonder how the reduced token output will affect the quality of responses.
Fatigue hype 2026 : le tri entre modèle et harness