288건 수집
2026-08-03 18:01

오늘의 핵심

DeepSeek-V4-플래시 업데이트

해커뉴스 (707포인트, 333댓글)

Hacker News

🤖 모델 & 제품 (9/23건)

OpenAI 2026-08-03

6개월 만에 반응형 음성 AI를 위한 실시간 시스템을 구축한 방법

GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.

OpenAI음성
OpenAI Blog
OpenAI 2026-08-01

수학과 이론 컴퓨터 과학의 10가지 발전

OpenAI는 기하학, 암호화 및 복잡성의 발전을 포함하여 수학과 이론 컴퓨터 과학의 오랫동안 공개된 문제에 대한 새로운 결과를 공유합니다.

OpenAI
OpenAI Blog
OpenAI 2026-07-31

Advancing responsible AI across Europe

OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will continue as the EU AI Act advances.

OpenAI규제안전
OpenAI Blog
OpenAI 2026-07-31

풍부한 지능 구축

A full-stack approach to making advanced AI more capable, more affordable, and more widely useful.

OpenAI Blog
OpenAI 2026-07-31

Univé는 AI 지원 인력을 구축합니다.

Univé가 리더십, 책임 있는 거버넌스, 직원 주도 혁신을 결합하여 대규모로 업무를 혁신함으로써 ChatGPT Enterprise를 통해 AI 지원 인력을 구축한 방법을 알아보세요.

규제
OpenAI Blog
OpenAI 2026-07-31

범죄적 사기 작전 방해

OpenAI disrupted a Cambodia-based scam operation using ChatGPT to support investment, romance, gambling, and impersonation schemes.

OpenAI트레이딩
OpenAI Blog
xAI 2026-07-31

참조가 포함된 비디오 1.5를 상상해 보세요.

Our best video model, now with text, image, and voice references — generating up to 1080p.

비전음성
xAI Blog
Anthropic 2026-07-30

Investigating three real-world incidents in our cybersecurity evaluations

In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below w

Claude트레이딩벤치마크
Anthropic Blog
Google 2026-07-28

Gemini API Managed Agents: 3.6 Flash, hooks, and more

Managed Agents Gemini 3.6 Flash, Hooks and Triggers

Google에이전트
Google AI Blog

🌎 업계 동향 (11/141건)

News 2026-08-04

After killer quarter, Palantir CEO Alex Karp calls AI industry ‘Marxist’

After a quarter that delivered $1 billion in profit, Palantir CEO Alex Karp on Monday once again warned that AI frontier labs are too untrustworthy for enterprises.

TechCrunch AI
News 2026-08-04

Qwen3.8-Max는 대담한 주장과 함께 출시됩니다. 에이전트 컴퓨터 사용에서 GPT-5.6 Sol Max 및 Fable 5보다 성능이 뛰어납니다.

중국의 전자상거래 및 클라우드 대기업인 Alibaba의 AI 연구원으로 구성된 Qwen 팀은 어젯밤 새로운 플래그십 2조 4천억 매개변수 전문가 혼합(MoE) 다중 모달 대형 언어 모델(LLM)인 Qwen3.8-Max를 공개했습니다. 이 모델은 첨단 AI 시장에서 가장 경쟁이 치열한 분야 중 하나인 자율 소프트웨어 엔지니어링 및 장기적 엔터프라이즈 작업을 목표로 합니다. 회사에서 벤치마킹을 발표한 경우

OpenAI에이전트연구벤치마크
VentureBeat AI
News 2026-08-04

Asana의 AI 에이전트는 회사 전체에서 메모리를 공유하지만 비밀은 공유하지 않습니다.

AI 에이전트를 구축하는 기업 팀은 계속 같은 벽에 부딪힙니다. 즉, 프롬프트에 응답할 수 있지만 마지막 5명이 요청한 내용을 기억할 수 없고 지난 달 버전이 실제로 작동했는지 여부를 알 수 없는 챗봇입니다. VB Transform 2026에서 VentureBeat의 Sam Witteveen과의 대화에서 Asana의 최고 제품 책임자인 Arnab Bose는 풀었습니다. 그의 팀이 이 문제를 어떻게 해결했는지

에이전트
VentureBeat AI
Community 2026-08-03

SQLite 중요 CVE 또는 LLM Slop?

The JFrog security research team recently identified a supply chain attack targeting the `xinference` package on PyPI. Versions 2.6.0, 2.6.1, and 2.6.2 were compromised and yanked by maintainers after users reported suspicious behavior. If you installed or imported these versions, you must assume yo

Hacker News · 695점 · 댓글 347
Community 2026-08-03

Prevent cognitive debt by manually retyping LLM-generated code

해커뉴스 (375포인트, 316댓글)

Hacker News · 375점 · 댓글 316
Community 2026-08-03

AI 생산성 격차

Why the productivity gains from AI are still small.

Hacker News · 104점 · 댓글 99
Community 2026-08-03

Show HN: Nightcrawler – 스마트폰에서 실행되는 로컬 AI 침투 테스트 에이전트

Local AI powered red teamer on a phone. Contribute to garagehq/nightcrawler development by creating an account on GitHub.

에이전트
Hacker News · 102점 · 댓글 30
Community 2026-08-03

What's the largest software project AI can complete on its own?

MirrorCode is Epoch AI's benchmark for long-horizon coding: AI can reimplement entire programs end-to-end, with no access to the original source code.

Hacker News · 66점 · 댓글 74
News 2026-08-03

유럽의 AI 라벨링 및 투명성 규칙이 이제 시행됩니다.

The European Union has ushered in some additional rules that aim to make it easier for people to identify chatbots and AI deepfakes online. The new transparency obligations under the bloc's landmark AI Act came into effect on August 2nd, requiring companies to disclose when people are interacti

The Verge AI
Community 2026-07-31

qm – 업무용 멀티플레이어 에이전트 하네스

업무용 멀티플레이어 에이전트 하네스. GitHub에 계정을 만들어 yc-software/qm 개발에 기여하세요.

에이전트
Hacker News · 650점 · 댓글 152
News 2026-07-31

Claude published malicious code to the Internet and attacked 3 real companies

해킹이 기존 방법을 사용했다면 누군가 감옥에 갔을 가능성이 높습니다.

Claude
Ars Technica AI

📚 논문 & 연구 (18/114건)

arXiv 2026-07-31

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided e

에이전트가격벤치마크
arXiv (cs.AI)
arXiv 2026-07-31

FDD-ON 개발: VAV HVAC 시스템 오류 감지 및 진단을 위한 온톨로지

Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency, and maintenance effectiveness. However, effective deployment of FDD solutions in buildings requires structured domain knowledge that can bridge heterogeneous data sources, diverse eq

arXiv (cs.AI)
arXiv 2026-07-31

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

LLM이 코드 완성 시스템에서 자율적인 과학 에이전트로 발전함에 따라 실험 수행 능력을 평가하는 것이 점점 더 중요해지고 있습니다. 기존 벤치마크는 일반적으로 정적 코드 생성, 문서 복제 또는 최종 답변 정확성에 중점을 두지만 에이전트가 실험적 증거를 해석하고 이를 사용하여 후속 하이퍼파라미터 dec을 안내할 수 있는지 여부를 직접 평가하지는 않습니다.

에이전트코딩연구벤치마크
arXiv (cs.AI)
arXiv 2026-07-31

The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

전통적인 정적 평가는 종종 야망에 불이익을 주고 진단 피드백을 모호하게 만드는 감산 적자 기반 채점 모델에 의존합니다. 반대로, 전통적인 대면 구술 시험은 학업 계층에 내재된 사회학적 권력 불균형과 수행 불안을 악화시킴으로써 구조와 무관한 심각한 차이를 도입합니다. 이 논문은 이론적 기초를 제시합니다.

연구
arXiv (cs.AI)
arXiv 2026-07-31

CENDRe: 자연 도메인 표현을 사용한 개념 추출

CNN(컨벌루션 신경망)은 시계열 분류에 널리 사용되지만 중요한 도메인에 배포하려면 예측을 주도하는 시간적 및 스펙트럼 패턴을 이해해야 합니다. 개념 추출(CE) 방법은 모델의 잠재 공간 내 표현을 분석하여 이러한 패턴을 식별합니다. 그러나 기존 시계열 CE 방법에는 세 가지 제한 사항이 있습니다.

arXiv (cs.AI)
arXiv 2026-07-31

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding errors and performance plateaus, particularly when th

에이전트로봇규제
arXiv (cs.AI)
arXiv 2026-07-31

Differentially Private Nonparametric Modal Learning with Applications to Regression and Clustering

Density modes provide a localized and interpretable summary of multimodal distributions, but their estimation under rigorous differential privacy constraints remains largely unexplored. We study differentially private recovery of density modes for multivariate distributions under local smoothness, c

arXiv (cs.LG)
arXiv 2026-07-31

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback

SignMuon compresses the Muon update to one bit per parameter by taking its elementwise sign, providing the most direct way to run a matrix-aware optimizer under an extremely low communication budget. It outperforms SignSGD in practice, yet it can ascend even on a linear function. Signing the gradien

arXiv (cs.LG)
arXiv 2026-07-31

동결 후 선택: 희소 관측에서 PDE 발견을 위한 구조화된 필드 어댑터 및 안정성이 검증된 약한 선택

PDE discovery from sparse observations requires reconstructing a continuous field and selecting the correct differential terms. Our analysis of optimization paths in coupled neural PDE discovery reveals three behaviors: the exact support can persist to the end of training, appear only transiently, o

arXiv (cs.LG)
arXiv 2026-07-31

GQ-FSL: 녹색 양자화 연합 분할 학습

Deploying state-of-the-art deep neural networks (DNNs) at the wireless edge is severely bottlenecked by the strict energy and resource constraints of mobile devices. While federated split learning (FSL) mitigates on-device computation by offloading workloads to an edge server, this may introduce sys

arXiv (cs.LG)
arXiv 2026-07-31

QASP: Query-Adaptive Robust Vector Search Policy

A fundamental challenge of vector search is achieving consistently high recall while minimizing computational costs. Fixed search parameters cause significant performance variance across queries, and conventional evaluation on average recall masks these per-query disparities. We introduce QASP (Quer

규제가격벤치마크
arXiv (cs.LG)
arXiv 2026-07-31

The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs

Parameter-Efficient Fine-Tuning (PEFT) commonly adapts large language models using a single shared Low-Rank Adapter (LoRA). This shared optimization space often suffers from interference when adapting heterogeneous task sequences, leading to poor transfer and catastrophic forgetting. Existing approa

연구규제
arXiv (cs.LG)
arXiv 2026-07-31

TokTier: Agentic LLM 서비스를 위한 정확한 상태 저장 토큰화

LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small tool result, and reuse is hard because even a short append can change token boundaries near the end

에이전트코딩가격
arXiv (cs.CL)
arXiv 2026-07-31

Evolving language compositionality in a frequency-structured meaning space

반복 학습 모델은 언어 진화를 조사하기 위해 도입되었습니다. 인간 언어의 특징적인 속성이 적어도 부분적으로 한 언어 사용자에서 다른 언어 사용자로의 반복적인 전송에 의해 형성되는 방식입니다. 핵심 발견은 언어가 언어 학습 병을 통해 반복적으로 전달된 결과로 언어 구성성이 자발적으로 발생할 수 있다는 것입니다.

트레이딩
arXiv (cs.CL)
arXiv 2026-07-31

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM backbone latents, which

로봇비전
arXiv (cs.CL)
arXiv 2026-07-31

ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression

KV 캐시 압축은 효율적인 장기 컨텍스트 추론을 위해 필수적입니다. 기존 퇴거 방법은 선택되지 않은 토큰을 영구적으로 폐기하고 결과적으로 관심에 대한 총 기여도를 제거합니다. 병합 기반 대안은 더 많은 정보를 보존하지만 정확하게 유지되어야 하는 유지된 키와 값을 교란시킬 수 있습니다. 우리는 캐시 제거에 의해 생략된 정보가 다음과 같이 공식화될 수 있음을 관찰했습니다.

arXiv (cs.CL)
arXiv 2026-07-31

아첨은 협력적인 비전-언어 작업에서 인식론적 경계를 약화시킵니다.

To maintain common ground in cooperative conversation, humans iteratively update their beliefs as conversation participants share new information; participants who are epistemically vigilant detect when new information conflicts with prior beliefs and take steps to repair these conflicts. In order f

비전
arXiv (cs.CL)
arXiv 2026-07-31

ARB: AI 텍스트 탐지기 평가를 위한 일치된 저자 재작성 벤치마크 데이터 세트

Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it remains unclear whether performance measured on this conventional benchmar

벤치마크
arXiv (cs.CL)

🎥 영상 & 튜토리얼 (3/10건)

YouTube 2026-08-03

또 다른 DeepSeek 순간이 왔습니다

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 DeepSeek v4 Flash 0731: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 DeepSeek API: https://platform.deepseek.com/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers poss

DeepSeek연구
Two Minute Papers
YouTube 2026-08-01

ChatGPT를 나만의 Jarvis로 전환하기

Here's how to turn AI into your own personal Jarvis 🤖 By using the voice features on ChatGPT and Claude, you can now build, edit, and iterate on interactive 3D designs in real-time—completely hands-free. I used voice commands to generate a 3D Iron Man suit, scale the helmet, change the colors,

Claude음성
Matt Wolfe AI
YouTube 2026-07-29

Anthropic이 방금 인디 해커를 죽였나요...?

무료로 Blacksmith를 사용해 GitHub Actions를 2배 더 빠르게 실행하세요. - https://www.blacksmith.sh/ Anthropic이 방금 Opus 5를 출시했고 인디 해커의 꿈이 무너질 수도 있습니다. 자세히 살펴보겠습니다. #coding #programming 더 많은 Fireship을 원하시나요? 🗞️ 뉴스레터: https://bytes.dev 🧠 강좌: https://fireship.dev

Anthropic코딩
Fireship