🤖 모델 & 제품 (9/23건)
6개월 만에 반응형 음성 AI를 위한 실시간 시스템을 구축한 방법
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
수학과 이론 컴퓨터 과학의 10가지 발전
OpenAI는 기하학, 암호화 및 복잡성의 발전을 포함하여 수학과 이론 컴퓨터 과학의 오랫동안 공개된 문제에 대한 새로운 결과를 공유합니다.
Advancing responsible AI across Europe
OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will continue as the EU AI Act advances.
풍부한 지능 구축
A full-stack approach to making advanced AI more capable, more affordable, and more widely useful.
Univé는 AI 지원 인력을 구축합니다.
Univé가 리더십, 책임 있는 거버넌스, 직원 주도 혁신을 결합하여 대규모로 업무를 혁신함으로써 ChatGPT Enterprise를 통해 AI 지원 인력을 구축한 방법을 알아보세요.
범죄적 사기 작전 방해
OpenAI disrupted a Cambodia-based scam operation using ChatGPT to support investment, romance, gambling, and impersonation schemes.
참조가 포함된 비디오 1.5를 상상해 보세요.
Our best video model, now with text, image, and voice references — generating up to 1080p.
Investigating three real-world incidents in our cybersecurity evaluations
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below w
Gemini API Managed Agents: 3.6 Flash, hooks, and more
Managed Agents Gemini 3.6 Flash, Hooks and Triggers
🌎 업계 동향 (11/141건)
After killer quarter, Palantir CEO Alex Karp calls AI industry ‘Marxist’
After a quarter that delivered $1 billion in profit, Palantir CEO Alex Karp on Monday once again warned that AI frontier labs are too untrustworthy for enterprises.
Qwen3.8-Max는 대담한 주장과 함께 출시됩니다. 에이전트 컴퓨터 사용에서 GPT-5.6 Sol Max 및 Fable 5보다 성능이 뛰어납니다.
중국의 전자상거래 및 클라우드 대기업인 Alibaba의 AI 연구원으로 구성된 Qwen 팀은 어젯밤 새로운 플래그십 2조 4천억 매개변수 전문가 혼합(MoE) 다중 모달 대형 언어 모델(LLM)인 Qwen3.8-Max를 공개했습니다. 이 모델은 첨단 AI 시장에서 가장 경쟁이 치열한 분야 중 하나인 자율 소프트웨어 엔지니어링 및 장기적 엔터프라이즈 작업을 목표로 합니다. 회사에서 벤치마킹을 발표한 경우
Asana의 AI 에이전트는 회사 전체에서 메모리를 공유하지만 비밀은 공유하지 않습니다.
AI 에이전트를 구축하는 기업 팀은 계속 같은 벽에 부딪힙니다. 즉, 프롬프트에 응답할 수 있지만 마지막 5명이 요청한 내용을 기억할 수 없고 지난 달 버전이 실제로 작동했는지 여부를 알 수 없는 챗봇입니다. VB Transform 2026에서 VentureBeat의 Sam Witteveen과의 대화에서 Asana의 최고 제품 책임자인 Arnab Bose는 풀었습니다. 그의 팀이 이 문제를 어떻게 해결했는지
SQLite 중요 CVE 또는 LLM Slop?
The JFrog security research team recently identified a supply chain attack targeting the `xinference` package on PyPI. Versions 2.6.0, 2.6.1, and 2.6.2 were compromised and yanked by maintainers after users reported suspicious behavior. If you installed or imported these versions, you must assume yo
Prevent cognitive debt by manually retyping LLM-generated code
해커뉴스 (375포인트, 316댓글)
AI 생산성 격차
Why the productivity gains from AI are still small.
Show HN: Nightcrawler – 스마트폰에서 실행되는 로컬 AI 침투 테스트 에이전트
Local AI powered red teamer on a phone. Contribute to garagehq/nightcrawler development by creating an account on GitHub.
What's the largest software project AI can complete on its own?
MirrorCode is Epoch AI's benchmark for long-horizon coding: AI can reimplement entire programs end-to-end, with no access to the original source code.
유럽의 AI 라벨링 및 투명성 규칙이 이제 시행됩니다.
The European Union has ushered in some additional rules that aim to make it easier for people to identify chatbots and AI deepfakes online. The new transparency obligations under the bloc's landmark AI Act came into effect on August 2nd, requiring companies to disclose when people are interacti
qm – 업무용 멀티플레이어 에이전트 하네스
업무용 멀티플레이어 에이전트 하네스. GitHub에 계정을 만들어 yc-software/qm 개발에 기여하세요.
Claude published malicious code to the Internet and attacked 3 real companies
해킹이 기존 방법을 사용했다면 누군가 감옥에 갔을 가능성이 높습니다.
📚 논문 & 연구 (18/114건)
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided e
FDD-ON 개발: VAV HVAC 시스템 오류 감지 및 진단을 위한 온톨로지
Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency, and maintenance effectiveness. However, effective deployment of FDD solutions in buildings requires structured domain knowledge that can bridge heterogeneous data sources, diverse eq
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers
LLM이 코드 완성 시스템에서 자율적인 과학 에이전트로 발전함에 따라 실험 수행 능력을 평가하는 것이 점점 더 중요해지고 있습니다. 기존 벤치마크는 일반적으로 정적 코드 생성, 문서 복제 또는 최종 답변 정확성에 중점을 두지만 에이전트가 실험적 증거를 해석하고 이를 사용하여 후속 하이퍼파라미터 dec을 안내할 수 있는지 여부를 직접 평가하지는 않습니다.
The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
전통적인 정적 평가는 종종 야망에 불이익을 주고 진단 피드백을 모호하게 만드는 감산 적자 기반 채점 모델에 의존합니다. 반대로, 전통적인 대면 구술 시험은 학업 계층에 내재된 사회학적 권력 불균형과 수행 불안을 악화시킴으로써 구조와 무관한 심각한 차이를 도입합니다. 이 논문은 이론적 기초를 제시합니다.
CENDRe: 자연 도메인 표현을 사용한 개념 추출
CNN(컨벌루션 신경망)은 시계열 분류에 널리 사용되지만 중요한 도메인에 배포하려면 예측을 주도하는 시간적 및 스펙트럼 패턴을 이해해야 합니다. 개념 추출(CE) 방법은 모델의 잠재 공간 내 표현을 분석하여 이러한 패턴을 식별합니다. 그러나 기존 시계열 CE 방법에는 세 가지 제한 사항이 있습니다.
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning
Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding errors and performance plateaus, particularly when th
Differentially Private Nonparametric Modal Learning with Applications to Regression and Clustering
Density modes provide a localized and interpretable summary of multimodal distributions, but their estimation under rigorous differential privacy constraints remains largely unexplored. We study differentially private recovery of density modes for multivariate distributions under local smoothness, c
Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback
SignMuon compresses the Muon update to one bit per parameter by taking its elementwise sign, providing the most direct way to run a matrix-aware optimizer under an extremely low communication budget. It outperforms SignSGD in practice, yet it can ascend even on a linear function. Signing the gradien
동결 후 선택: 희소 관측에서 PDE 발견을 위한 구조화된 필드 어댑터 및 안정성이 검증된 약한 선택
PDE discovery from sparse observations requires reconstructing a continuous field and selecting the correct differential terms. Our analysis of optimization paths in coupled neural PDE discovery reveals three behaviors: the exact support can persist to the end of training, appear only transiently, o
GQ-FSL: 녹색 양자화 연합 분할 학습
Deploying state-of-the-art deep neural networks (DNNs) at the wireless edge is severely bottlenecked by the strict energy and resource constraints of mobile devices. While federated split learning (FSL) mitigates on-device computation by offloading workloads to an edge server, this may introduce sys
QASP: Query-Adaptive Robust Vector Search Policy
A fundamental challenge of vector search is achieving consistently high recall while minimizing computational costs. Fixed search parameters cause significant performance variance across queries, and conventional evaluation on average recall masks these per-query disparities. We introduce QASP (Quer
The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs
Parameter-Efficient Fine-Tuning (PEFT) commonly adapts large language models using a single shared Low-Rank Adapter (LoRA). This shared optimization space often suffers from interference when adapting heterogeneous task sequences, leading to poor transfer and catastrophic forgetting. Existing approa
TokTier: Agentic LLM 서비스를 위한 정확한 상태 저장 토큰화
LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small tool result, and reuse is hard because even a short append can change token boundaries near the end
Evolving language compositionality in a frequency-structured meaning space
반복 학습 모델은 언어 진화를 조사하기 위해 도입되었습니다. 인간 언어의 특징적인 속성이 적어도 부분적으로 한 언어 사용자에서 다른 언어 사용자로의 반복적인 전송에 의해 형성되는 방식입니다. 핵심 발견은 언어가 언어 학습 병을 통해 반복적으로 전달된 결과로 언어 구성성이 자발적으로 발생할 수 있다는 것입니다.
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM backbone latents, which
ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression
KV 캐시 압축은 효율적인 장기 컨텍스트 추론을 위해 필수적입니다. 기존 퇴거 방법은 선택되지 않은 토큰을 영구적으로 폐기하고 결과적으로 관심에 대한 총 기여도를 제거합니다. 병합 기반 대안은 더 많은 정보를 보존하지만 정확하게 유지되어야 하는 유지된 키와 값을 교란시킬 수 있습니다. 우리는 캐시 제거에 의해 생략된 정보가 다음과 같이 공식화될 수 있음을 관찰했습니다.
아첨은 협력적인 비전-언어 작업에서 인식론적 경계를 약화시킵니다.
To maintain common ground in cooperative conversation, humans iteratively update their beliefs as conversation participants share new information; participants who are epistemically vigilant detect when new information conflicts with prior beliefs and take steps to repair these conflicts. In order f
ARB: AI 텍스트 탐지기 평가를 위한 일치된 저자 재작성 벤치마크 데이터 세트
Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it remains unclear whether performance measured on this conventional benchmar
🎥 영상 & 튜토리얼 (3/10건)
또 다른 DeepSeek 순간이 왔습니다
❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 DeepSeek v4 Flash 0731: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 DeepSeek API: https://platform.deepseek.com/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers poss
ChatGPT를 나만의 Jarvis로 전환하기
Here's how to turn AI into your own personal Jarvis 🤖 By using the voice features on ChatGPT and Claude, you can now build, edit, and iterate on interactive 3D designs in real-time—completely hands-free. I used voice commands to generate a 3D Iron Man suit, scale the helmet, change the colors,
Anthropic이 방금 인디 해커를 죽였나요...?
무료로 Blacksmith를 사용해 GitHub Actions를 2배 더 빠르게 실행하세요. - https://www.blacksmith.sh/ Anthropic이 방금 Opus 5를 출시했고 인디 해커의 꿈이 무너질 수도 있습니다. 자세히 살펴보겠습니다. #coding #programming 더 많은 Fireship을 원하시나요? 🗞️ 뉴스레터: https://bytes.dev 🧠 강좌: https://fireship.dev