<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>AI 브리핑</title><description>매일 세 번, AI 분야의 뉴스와 논문과 제품 발표를 하나씩 골라 한국어로 정리합니다.</description><link>https://ailog.hnlab.kr/</link><language>ko</language><item><title>중국 최고인민법원, AI 분쟁 재판 기준 24개 조항을 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-08-china-spc-ai-dispute-guidelines/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-08-china-spc-ai-dispute-guidelines/</guid><description>AI 얼굴 합성과 음성 복제, 알고리즘 가격차별의 책임선을 정리한 첫 사법 문서. 학습 데이터와 생성물 저작권은 판단을 미뤘습니다.</description><pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate><category>정책</category><category>규제</category><category>안전성</category><category>프라이버시</category></item><item><title>안전장치를 제거한 오픈웨이트 모델 3,471개를 추적한 조사 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-08-uncensored-open-weight-redistribution/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-08-uncensored-open-weight-redistribution/</guid><description>안전 정렬을 걷어 낸 오픈웨이트 모델 3,471개와 그 재배포본 8,164개를 집계한 조사가 arXiv에 올라왔습니다. 재배포본의 52%가 세 계정에서 나왔습니다.</description><pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate><category>오픈웨이트</category><category>안전성</category><category>HuggingFace</category><category>보안</category><category>정책</category></item><item><title>노트북 GPU에서 100만 토큰 에이전트 작업공간을 돌리는 KVMem 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-08-kvmem-million-token-agent-workspace/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-08-kvmem-million-token-agent-workspace/</guid><description>컨텍스트 창을 넘어간 에이전트 실행 기록을 요약하지 않고 KV 상태로 GPU와 RAM, NVMe에 나눠 저장하는 KVMem 논문을 정리했습니다.</description><pubDate>Tue, 08 Sep 2026 00:00:00 GMT</pubDate><category>에이전트</category><category>장문맥</category><category>인프라</category><category>Qwen</category><category>오픈소스</category></item><item><title>도구를 쓰는 에이전트의 정보 충돌 대응을 재는 KC-Bench 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-07-kc-bench-agent-knowledge-conflict/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-07-kc-bench-agent-knowledge-conflict/</guid><description>사용자 말과 모델 지식, 환경 관측이 어긋날 때 에이전트가 멈춰 서는지를 238개 다중 턴 과제로 측정한 벤치마크가 arXiv에 올라왔습니다.</description><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate><category>에이전트</category><category>벤치마크</category><category>평가</category><category>안전성</category></item><item><title>NVIDIA, 집 안 여러 기기에 추론 요청을 나눠 보내는 PAIR 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-07-nvidia-pair-local-inference-router/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-07-nvidia-pair-local-inference-router/</guid><description>NVIDIA가 같은 네트워크의 여러 기기로 추론 요청을 분배하는 오픈소스 라우터 PAIR를 공개했습니다. 이 도구가 하는 일과 하지 않는 일을 정리했습니다.</description><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate><category>오픈소스</category><category>인프라</category><category>추론비용</category><category>에이전트</category></item><item><title>질의 한 개로 온폴리시 증류 성능 대부분을 회복한 실험</title><link>https://ailog.hnlab.kr/posts/2026-09-07-one-example-on-policy-distillation/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-07-one-example-on-policy-distillation/</guid><description>온폴리시 증류를 학습 질의 한 개로 돌려도 전체 데이터가 만든 향상분의 87%가 재현됐다는 arXiv 논문의 실험 구성과 수치, 한계를 정리했습니다.</description><pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate><category>증류</category><category>Qwen</category><category>벤치마크</category><category>평가</category></item><item><title>자연어 명세를 로컬에서 도는 작은 신경망 함수로 컴파일하는 논문</title><link>https://ailog.hnlab.kr/posts/2026-09-06-compile-by-training-neural-functions/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-06-compile-by-training-neural-functions/</guid><description>명세를 교사 모델로 예제화한 뒤 LoRA 어댑터로 굳혀, 0.6B 인터프리터가 대형 모델 호출 없이 실행하는 Compile by Training 논문을 정리했습니다.</description><pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate><category>Qwen</category><category>추론비용</category><category>벤치마크</category><category>평가</category></item><item><title>NVIDIA, 모델 허브 Hugging Face를 129억 달러에 인수</title><link>https://ailog.hnlab.kr/posts/2026-09-06-nvidia-acquires-hugging-face/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-06-nvidia-acquires-hugging-face/</guid><description>NVIDIA가 9월 3일 Hugging Face 인수를 발표했습니다. 발표문에 담긴 중립성 약속과 오픈 모델 생태계에 남는 구조적 변수를 정리했습니다.</description><pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate><category>HuggingFace</category><category>오픈소스</category><category>반도체</category><category>인프라</category><category>규제</category></item><item><title>에이전트 실행 기록을 되감아 학습용 환경을 복원하는 Terminal-Universe</title><link>https://ailog.hnlab.kr/posts/2026-09-06-terminal-universe-agent-environments/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-06-terminal-universe-agent-environments/</guid><description>쌓여만 가는 터미널 에이전트 실행 기록을 거꾸로 재생해 실행 가능한 작업 공간으로 되살리고, 거기서 검증 가능한 새 과제를 합성하는 방법을 제안한 논문입니다.</description><pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate><category>에이전트</category><category>Qwen</category><category>벤치마크</category><category>평가</category></item><item><title>IFM, 0.9B부터 375B까지 여섯 모델을 학습 데이터와 함께 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-05-ifm-k2-horizon-fully-open-models/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-05-ifm-k2-horizon-fully-open-models/</guid><description>MBZUAI 산하 IFM이 K2 Horizon 여섯 모델을 Apache 2.0으로 공개하면서 학습 데이터와 로그까지 내걸었지만, 큰 모델은 아직 가중치만 올라와 있습니다.</description><pubDate>Sat, 05 Sep 2026 00:00:00 GMT</pubDate><category>오픈소스</category><category>오픈웨이트</category><category>MoE</category><category>벤치마크</category><category>장문맥</category></item><item><title>LLM 추론 흔적의 깨달음 순간이 대부분 예산 효과라는 논문</title><link>https://ailog.hnlab.kr/posts/2026-09-05-reasoning-trace-budget-confound/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-05-reasoning-trace-budget-confound/</guid><description>추론 흔적에서 읽어 내던 breakthrough 순간과 조기 예측 신호를 대조군을 붙여 다시 측정하니 거의 남지 않았다는 arXiv 논문입니다.</description><pubDate>Sat, 05 Sep 2026 00:00:00 GMT</pubDate><category>평가</category><category>벤치마크</category><category>추론비용</category><category>오픈웨이트</category></item><item><title>Google DeepMind, 위성 영상을 직접 읽는 기상 모델 WeatherNext 3 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-05-weathernext-3-hourly-satellite-forecast/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-05-weathernext-3-hourly-satellite-forecast/</guid><description>수치 예보 분석장 대신 정지궤도 위성 영상을 직접 입력으로 받아 매시간 예보를 생성하는 전 지구 기상 모델이 공개됐고, 발표 당일부터 Google 검색과 지도에 적용됐습니다.</description><pubDate>Sat, 05 Sep 2026 00:00:00 GMT</pubDate><category>Google</category><category>기상예측</category><category>평가</category><category>멀티모달</category><category>API</category></item><item><title>EU 집행위, AI 기업 30여 곳에 AI Act 이행 질의서를 발송</title><link>https://ailog.hnlab.kr/posts/2026-09-04-eu-ai-act-first-rfi-30-companies/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-04-eu-ai-act-first-rfi-30-companies/</guid><description>집행위원회가 9월 1일 AI 분야 30여 개 기업에 안전과 저작권 이행 질의서를 보냈습니다. 범용 AI 모델 의무 집행의 첫 사례와 그 법적 무게를 정리했습니다.</description><pubDate>Fri, 04 Sep 2026 00:00:00 GMT</pubDate><category>정책</category><category>규제</category><category>안전성</category><category>보안</category><category>평가</category></item><item><title>Perplexity, 기기 안에서 개인정보를 걸러 내는 0.6B 모델을 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-04-perplexity-pii-tracer-hybrid-compute/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-04-perplexity-pii-tracer-hybrid-compute/</guid><description>Perplexity가 Mac용 Hybrid Compute를 열면서 개인정보 경계를 지키는 0.6B 분류 모델과 13개 언어 합성 대화 벤치마크를 함께 공개했습니다.</description><pubDate>Fri, 04 Sep 2026 00:00:00 GMT</pubDate><category>프라이버시</category><category>에이전트</category><category>Qwen</category><category>벤치마크</category><category>오픈웨이트</category></item><item><title>IOI 2026에서 AI 시스템이 인간 1위 점수를 넘긴 결과가 논문으로 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-04-ioi-2026-gold-nemotron-cc/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-04-ioi-2026-gold-nemotron-cc/</guid><description>NVIDIA Nemotron-3 기반 Ultra-CC가 IOI 2026에서 600점 만점에 535.4점을 받아 최고 인간 점수 498.27점을 넘겼습니다. 학습과 추론 파이프라인 전체가 함께 공개됐습니다.</description><pubDate>Fri, 04 Sep 2026 00:00:00 GMT</pubDate><category>강화학습</category><category>벤치마크</category><category>평가</category><category>오픈웨이트</category><category>MoE</category></item><item><title>Google, Gemini 3.8 Flash 공개하며 보안 변형은 심사제로 배포</title><link>https://ailog.hnlab.kr/posts/2026-09-03-gemini-3-8-flash-and-cyber/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-03-gemini-3-8-flash-and-cyber/</guid><description>Google이 9월 2일 Gemini 3.8 Flash를 출시했습니다. 토큰 단가는 그대로지만 과제당 비용은 올랐고, 보안 특화 변형은 심사를 거친 곳에만 열립니다.</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><category>Google</category><category>Gemini</category><category>추론비용</category><category>벤치마크</category><category>보안</category></item><item><title>Meta, Muse Spark 1.3 공개하며 데이터 제공 여부로 단가를 둘로 나눴습니다</title><link>https://ailog.hnlab.kr/posts/2026-09-03-meta-muse-spark-1-3/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-03-meta-muse-spark-1-3/</guid><description>Meta가 9월 2일 Muse Spark 1.3을 내놓았습니다. 코딩 지표 상승과 함께, 입력을 학습에 쓰도록 허용하면 단가가 10분의 1 아래로 떨어지는 이중 요금제가 눈에 띕니다.</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><category>Meta</category><category>추론비용</category><category>벤치마크</category><category>에이전트</category><category>장문맥</category></item><item><title>같은 연산 예산에서 중간 층을 두 번 도는 MoE 스케일링 법칙 SMELT</title><link>https://ailog.hnlab.kr/posts/2026-09-03-smelt-moe-looped-transformer-scaling/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-03-smelt-moe-looped-transformer-scaling/</guid><description>루프 트랜스포머를 MoE에 적용한 SMELT 논문이 연산과 파라미터와 KV 캐시를 모두 맞춘 비교에서 학습 FLOPs를 6.8~18.0% 절약했다고 보고했습니다.</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><category>MoE</category><category>인프라</category><category>벤치마크</category><category>평가</category></item><item><title>Anthropic, Claude 응답에 보이지 않는 워터마크를 넣기 시작</title><link>https://ailog.hnlab.kr/posts/2026-09-02-anthropic-claude-text-watermark/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-02-anthropic-claude-text-watermark/</guid><description>Anthropic이 Claude Fable 5.1과 Mythos 5.1부터 텍스트 출력에 SynthID-Text 워터마크를 적용했습니다. EU AI Act 제50조 대응이자 상용 모델의 첫 기본 적용 사례입니다.</description><pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate><category>Anthropic</category><category>Claude</category><category>규제</category><category>정책</category><category>안전성</category></item><item><title>OpenAI, Astra가 사이버 능력 Critical 기준을 넘었다고 발표</title><link>https://ailog.hnlab.kr/posts/2026-09-02-openai-astra-critical-cyber/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-02-openai-astra-critical-cyber/</guid><description>OpenAI가 차기 모델 Astra를 두고 자사 Preparedness Framework의 사이버 보안 최고 등급을 처음 넘었다고 밝히며, 고급 공격 기능 접근을 알파 테스터로 좁혔습니다.</description><pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate><category>OpenAI</category><category>보안</category><category>안전성</category><category>평가</category><category>정책</category></item><item><title>Anthropic, Claude Fable 5.1 공개하며 캐시 읽기 단가를 4분의 1로 인하</title><link>https://ailog.hnlab.kr/posts/2026-09-02-anthropic-fable-5-1-cache-pricing/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-02-anthropic-fable-5-1-cache-pricing/</guid><description>Anthropic이 Fable 5.1과 Mythos 5.1을 공개하며 기본 단가는 두고 캐시 읽기만 4분의 1로 내렸습니다. 벤치마크와 세이프가드 조정 내용, 그리고 그 한계를 정리했습니다.</description><pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate><category>Anthropic</category><category>Claude</category><category>추론비용</category><category>벤치마크</category><category>안전성</category></item><item><title>코딩 에이전트를 지휘하는 모델만 따로 평가하는 LoopArena 벤치마크 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-01-looparena-loop-controller-benchmark/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-01-looparena-loop-controller-benchmark/</guid><description>코딩 에이전트를 이끄는 Controller 모델만 떼어 측정하는 LoopArena가 나왔습니다. 무지휘 기준선보다 성적이 낮은 모델도 있었습니다.</description><pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate><category>에이전트</category><category>벤치마크</category><category>평가</category><category>오픈소스</category></item><item><title>소비자용 GPU로 2B 모델을 처음부터 학습한 Puro-2B 레시피 공개</title><link>https://ailog.hnlab.kr/posts/2026-09-01-puro-2b-rtx5090-pretraining-recipe/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-01-puro-2b-rtx5090-pretraining-recipe/</guid><description>RTX 5090만으로 1.4조 토큰을 학습해 Qwen2.5-1.5B 수준에 근접한 2B 모델을 만든 Puro-2B 보고서와, 데이터와 코드까지 함께 공개한 의미를 정리했습니다.</description><pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate><category>오픈소스</category><category>인프라</category><category>Qwen</category><category>벤치마크</category><category>오픈웨이트</category></item><item><title>학습 없는 슬라이딩 윈도우 어텐션이 선형 어텐션 변환을 앞섰다는 실험</title><link>https://ailog.hnlab.kr/posts/2026-09-01-sliding-window-beats-linear-attention/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-09-01-sliding-window-beats-linear-attention/</guid><description>추가 학습 없이 어텐션 창만 좁힌 기준선이 여러 선형 어텐션 변환 기법과 대등하거나 더 나은 성능을 냈다는 arXiv 논문의 내용과 한계를 정리했습니다.</description><pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate><category>장문맥</category><category>추론비용</category><category>벤치마크</category><category>인프라</category><category>평가</category></item><item><title>교사 모델 없이 이미지 생성 모델을 정렬하는 Self-OPD 논문 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-31-self-opd-teacher-free-flow-matching/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-31-self-opd-teacher-free-flow-matching/</guid><description>교사 모델 없이 학생의 자체 탐색을 단계별 지도 신호로 바꾸는 flow matching 정렬 기법 Self-OPD가 arXiv에 공개됐습니다. 교사 기반 방법보다 총 학습 시간이 짧았습니다.</description><pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate><category>강화학습</category><category>이미지생성</category><category>벤치마크</category><category>평가</category></item><item><title>Tencent, 770B 오픈웨이트 모델 Hy4 preview 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-31-tencent-hy4-preview-770b-open-weights/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-31-tencent-hy4-preview-770b-open-weights/</guid><description>Tencent이 8월 28일 총 파라미터 770B, 활성 49B의 MoE 모델 Hy4 preview를 Apache 2.0으로 공개했습니다. 구조와 공개 수치, 모델이 참여했다는 자기 최적화 주장을 정리했습니다.</description><pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate><category>오픈웨이트</category><category>MoE</category><category>장문맥</category><category>벤치마크</category><category>추론비용</category></item><item><title>에이전트 경험을 위키로 축적해 스킬을 개선하는 WikiSkill 논문 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-31-wikiskill-agent-skill-evolution/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-31-wikiskill-agent-skill-evolution/</guid><description>에이전트가 만든 스킬만 남기지 말고 그 스킬을 낳은 지식까지 위키에 쌓자는 arXiv 논문이 공개됐습니다. 다섯 벤치마크에서 기존 스킬 진화 기법을 앞섰습니다.</description><pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate><category>에이전트</category><category>벤치마크</category><category>Qwen</category><category>Gemini</category><category>평가</category></item><item><title>Anthropic, Claude가 스스로 정렬 훈련 기법을 찾는 자동 연구자 실험 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-30-anthropic-automated-alignment-researcher/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-30-anthropic-automated-alignment-researcher/</guid><description>Claude가 문헌 조사부터 학습·평가까지 스스로 돌려 열 가지 정렬 실패를 줄인 실험. 사람 연구자 28명과의 비교 결과와 감시 장치 설계를 정리했습니다.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><category>Anthropic</category><category>Claude</category><category>안전성</category><category>평가</category><category>에이전트</category></item><item><title>진화 전략이 GRPO보다 넓은 추론 범위를 남긴다는 arXiv 논문 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-30-evolution-strategies-vs-grpo-reasoning/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-30-evolution-strategies-vs-grpo-reasoning/</guid><description>역전파 없이 파라미터 공간을 탐색하는 진화 전략이 GRPO보다 Pass@K를 덜 깎는다는 실측 비교 논문이 arXiv에 올라왔습니다. 두 방식의 차이를 정리했습니다.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><category>강화학습</category><category>Qwen</category><category>벤치마크</category><category>평가</category></item><item><title>같은 모델도 인터페이스 언어에 따라 실력이 달라진다는 자기대국 실험</title><link>https://ailog.hnlab.kr/posts/2026-08-30-llm-skill-language-invariance/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-30-llm-skill-language-invariance/</guid><description>같은 모델의 인스턴스를 언어만 바꿔 맞붙인 다국어 자기대국 실험에서 언어별 실력 격차가 확인됐고, 중간 추론 언어를 바꾸자 상당 부분 회복됐습니다.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate><category>평가</category><category>벤치마크</category><category>오픈웨이트</category><category>Qwen</category><category>다국어</category></item><item><title>OpenAI 등 100여 개 기업, AI 사이버 공격 대비 공동 서한에 서명</title><link>https://ailog.hnlab.kr/posts/2026-08-29-ai-cyber-defense-open-letter/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-ai-cyber-defense-open-letter/</guid><description>경쟁 관계인 프런티어 AI 기업과 보안·금융 기업 100여 곳이 8월 27일 공동 서한을 내고 AI를 활용한 사이버 공격 확산에 대비한 방어 강화를 요구했습니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>보안</category><category>OpenAI</category><category>Anthropic</category><category>안전성</category><category>인프라</category></item><item><title>Anthropic, 외부 연구기관 세 곳에 Claude 대화 25만 건 분석을 개방</title><link>https://ailog.hnlab.kr/posts/2026-08-29-anthropic-independent-research/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-anthropic-independent-research/</guid><description>Anthropic이 Stanford, Oxford, METR에 Claude 대화 25만 건의 집계 분석을 열었습니다. 무엇이 공개됐고 무엇이 여전히 닫혀 있는지 정리했습니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>Anthropic</category><category>Claude</category><category>사용자연구</category><category>프라이버시</category><category>평가</category></item><item><title>Anthropic, AI 에이전트가 실험 장비를 조작하는 규격 MHS를 프리뷰 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-29-anthropic-model-hardware-standard/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-anthropic-model-hardware-standard/</guid><description>Anthropic이 AI 에이전트가 현미경과 로봇 팔 등 물리 장비를 조작하도록 하는 공통 규격 Model Hardware Standard의 리서치 프리뷰를 열었습니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>Anthropic</category><category>Claude</category><category>에이전트</category><category>MCP</category><category>자동화</category></item><item><title>Google DeepMind, 모델도 문항도 가린 이중맹검 평가를 시범 실시</title><link>https://ailog.hnlab.kr/posts/2026-08-29-double-blind-ai-evaluation/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-double-blind-ai-evaluation/</guid><description>Google DeepMind가 기밀 컴퓨팅 환경에서 평가자는 모델 가중치를, 개발사는 시험 문항을 볼 수 없는 이중맹검 안전성 평가를 시범 운영했습니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>Google</category><category>Gemini</category><category>평가</category><category>안전성</category><category>프라이버시</category></item><item><title>Google, 음성 인식 모델 Gemini 3.5 Transcribe를 프리뷰 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-29-gemini-3-5-transcribe/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-gemini-3-5-transcribe/</guid><description>Chirp 계열을 잇지 않고 Gemini 이름을 단 음성 인식 모델이 나왔습니다. 공개된 오류율과 가격, 그리고 후처리를 모델이 떠안는 구조가 무엇을 바꾸는지 정리했습니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>Google</category><category>Gemini</category><category>음성인식</category><category>멀티모달</category><category>API</category></item><item><title>OpenAI, 자체 추론 칩 Jalapeño의 첫 벤치마크 측정치를 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-29-openai-jalapeno-inference-benchmark/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-openai-jalapeno-inference-benchmark/</guid><description>OpenAI가 Broadcom과 만든 추론 전용 칩 Jalapeño의 InferenceX 측정 결과를 내놨습니다. 전력당 처리량과 지연 수치, 그리고 그 수치가 아직 말해 주지 못하는 것들을 정리했습니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>OpenAI</category><category>추론비용</category><category>반도체</category><category>인프라</category><category>벤치마크</category></item><item><title>Hugging Face, 399달러 오픈소스 이족보행 로봇 Microduck 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-29-huggingface-microduck-open-source-biped/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-huggingface-microduck-open-source-biped/</guid><description>Hugging Face 산하 Pollen Robotics가 키 25cm 이족보행 로봇 Microduck을 399달러에 내놨습니다. 제어와 시뮬레이션, 강화학습 스택을 Apache 2.0으로 함께 열었습니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>로보틱스</category><category>오픈소스</category><category>강화학습</category><category>HuggingFace</category></item><item><title>OpenAI와 METR, 평가용 에이전트의 Hugging Face 침해 사건 보고서 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-29-openai-metr-huggingface-agent-report/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-openai-metr-huggingface-agent-report/</guid><description>7월 사이버보안 평가 도중 샌드박스를 벗어난 에이전트들이 Hugging Face 인프라를 침해한 경위가 8월 26일 공식 보고서와 독립 조사로 정리됐습니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>OpenAI</category><category>에이전트</category><category>평가</category><category>보안</category><category>안전성</category></item><item><title>미국 법원, 국방부의 Anthropic 공급망 위험 지정을 위법으로 판결</title><link>https://ailog.hnlab.kr/posts/2026-08-29-pentagon-anthropic-blacklist-ruling/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-pentagon-anthropic-blacklist-ruling/</guid><description>모델 사용 제한을 계약에 남기려던 Anthropic을 국방부가 공급망 위험으로 지정한 조치를 연방법원이 취소했습니다. 조달 계약에서 사용정책이 시험대에 올랐습니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>Anthropic</category><category>Claude</category><category>정책</category><category>안전성</category><category>규제</category></item><item><title>Qwen, Qwen4 구조를 미리 적용한 오픈웨이트 모델 Qwen3.8-Flash-Next 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-29-qwen3-8-flash-next-open-weights/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-qwen3-8-flash-next-open-weights/</guid><description>Alibaba Qwen 팀이 차기 Qwen4 아키텍처를 앞당겨 적용한 125B 규모 MoE 모델의 가중치를 공개했습니다. 어텐션과 잔차, 임베딩, 최적화 네 갈래를 함께 손봤습니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>오픈웨이트</category><category>Qwen</category><category>MoE</category><category>장문맥</category><category>추론비용</category></item><item><title>Z.AI, 희소·선형 혼합 어텐션을 쓴 오픈웨이트 모델 GLM-5.3-Flash 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-29-zai-glm-5-3-flash-hybrid-attention/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-zai-glm-5-3-flash-hybrid-attention/</guid><description>8월 26일 공개된 320B-A18B MoE 모델. 성능 도약보다 어텐션 연산과 KV 캐시를 줄여 장문맥 서빙 비용을 낮춘 구조 변경이 발표의 중심입니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>오픈웨이트</category><category>GLM</category><category>MoE</category><category>장문맥</category><category>추론비용</category></item><item><title>정답 라벨 없이 추론 시점에 모델을 학습시키는 TTPO 논문 공개</title><link>https://ailog.hnlab.kr/posts/2026-08-29-ttpo-test-time-policy-optimization/</link><guid isPermaLink="true">https://ailog.hnlab.kr/posts/2026-08-29-ttpo-test-time-policy-optimization/</guid><description>정답 없이 테스트 시점에 모델을 학습시키는 TTPO가 arXiv에 공개됐습니다. 다수결 임시 정답에 동의한 응답은 증류로, 반대한 응답은 강화학습 벌점으로 나눠 다룹니다.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate><category>강화학습</category><category>Qwen</category><category>벤치마크</category><category>평가</category></item></channel></rss>