- 83df861 Internal change by Google AI Edge · 4 hours ago upstream/main
- b319bde Clean up `runner_test_suite.h` by Quentin Khan · 7 hours ago
- 31c0eca Allow the YNNPACK delegate to rebuild its subgraph outside of delegation without invoke GetNodeAndRegistration directly by Google AI Edge · 8 hours ago
- fe837be Support Gemma 4 26B INT4 MoE convert-path model on Metal GPU delegate. by Fengwu Yao · 8 hours ago
- bbb1312 Rename `LiteRtStaticLinkedAcceleratorCpuDef` to `LiteRtStaticLinkedAcceleratorXnnpackDef` and remove `__attribute__((weak))` from `RegisterCpuAccelerator`. by Gerardo Carranza · 11 hours ago
- 4c60cf4 Update CODEOWNERS file to include a url link pointing to github team by Google AI Edge · 11 hours ago
- 4018aed Merge pull request #10448 from graham0824:dev/hungjuiw/fix-addn-test2 by Copybara-Service · 11 hours ago
- 1052937 Internal change by Maria Lyubimtseva · 11 hours ago
- b0e378a Add a LiteRT cpu_backend build flag and a minimal build guide by Terry Heo · 11 hours ago
- dd07838 Implement WebGPU shape/indexing, spatial conv/pool, and LSTM operations and enable GPU NumericalTestSuite. by Ping Yu · 12 hours ago
- 97fe163 Merge pull request #10270 from gunes-arm:pr/naming-changes by Copybara-Service · 12 hours ago
- 0dedf30 Expose delegation metrics through LiteRT C and C++ compiled model APIs. by Gerardo Carranza · 13 hours ago
- 9b04b48 Remove obsolete highwayhash patching from Windows wheel build script by Terry Heo · 14 hours ago
- 42c8fe8 Fix -Wpass-failed build break in the XNNPACK MoE kernel under sanitizers. by Fengwu Yao · 14 hours ago
- eb610ee Use hardware simd_sum and single-phase shared-memory reduction in gated_delta_update shader. by Google AI Edge · 14 hours ago
- e52511b Make the OpenVINO android_x86_64 TAP project actually compile src_gen. by Matt Kreileder CA · 14 hours ago
- 5b39c03 PR #9968: Qualcomm AI Engine Direct - Switch from coarse-grained to fine-grained ops. by Jiun Kai Yang · 15 hours ago
- d703c55 Simplify GQA graph construction in YNNPACK and tighten SDPA decode1 threshold. by Volodymyr Kysenko · 15 hours ago
- 261401c Add a migration guide from TFLite Support to LiteRT Support. by Jun Jiang · 15 hours ago
- 2db38fd Optimize XNNPACK MoE expert weight dequantization, single-token GEMV, and scratch memory. by Fengwu Yao · 16 hours ago
- 9de33b5 Avoid linking XNNPACK in BuiltinOpResolverWithoutDefaultDelegates by Terry Heo · 17 hours ago
- 2f17ea5 Add YNNPACK CPU backend support to LiteRT ATS. by Gerardo Carranza · 17 hours ago
- ecdc41c Legalize REDUCE_ANY in the OpenVINO compiler plugin. by Tommy Chiang · 17 hours ago
- 724a247 Decouple YNNPACK CPU accelerator registration from compile-time preprocessor macros. by Gerardo Carranza · 17 hours ago
- 2134610 Add filegroup target for LiteRT C++ SDK sources by Catalin Termure · 18 hours ago
- 06705e9 Add mixed-precision and FP32 recurrent state support to LiteRT GPU delegate and custom ops. by Google AI Edge · 19 hours ago
- 541ebbb Move MakeConvWithBatchIds to common MoE utils and reuse it in LiteRT composite kernel. by Raman Sarokin · 19 hours ago
- 471a9fd Update README.md for LiteRT Swift user guide. by Jun Jiang · 27 hours ago
- cde1321 Enable BOOL mask pruning for all Flash-Decode head dims in sdpa_transposed. by Fengwu Yao · 30 hours ago
- dcfd9b7 Do not clamp Flash-Decode active tokens to param[0] for single-query causal SDPA. by Fengwu Yao · 31 hours ago
- 0d4a8cb Qualcomm AI Engine Direct - Fix addn_test.cc issue due to namespace by hungjuiw · 31 hours ago
- eed62af Support the tanh-approximated GELU activation (GeGLU) in the odml.swiglu GPU kernel. by Fengwu Yao · 31 hours ago
- b5de9a0 Support V RMSNorm, Q-only mode, partial RoPE proportion, and head_dim up to 512 in qkv_norm_rope. by Fengwu Yao · 32 hours ago
- 6c6af01 Remove the unused BUILD_CONVERTER option from the LiteRT wheel build by Terry Heo · 32 hours ago
- 7116bc9 Fix weight sharing, non-sharing compilation now produces separate context bin. by Andrew Zhang · 32 hours ago
- b50f4ba Update copyright header for colabs by Maria Lyubimtseva · 32 hours ago
- ef7ea74 Add the documentation and tests of nvidia's benchmark script by Google AI Edge · 33 hours ago
- 7970321 The NVIDIA dispatch library copied a CUDA tensor buffer to the device every by Google AI Edge · 33 hours ago
- 948486f Overlap the softmax and the score products of the tiled attention kernel by Google AI Edge · 33 hours ago
- 9a1d7f2 Run Gemma 4 prefill fully connected layers on the tensor-core GEMM by Google AI Edge · 33 hours ago
- 4122e7f Add a tensor-core GEMM for BF16 activations and INT4 weights by Google AI Edge · 33 hours ago
- 71684ba Accumulate the scores of BF16 queries in FP16 on the tiled attention kernel by Google AI Edge · 33 hours ago
- cffedb5 Run Gemma 4 prefill attention on a tiled tensor-core kernel by Google AI Edge · 33 hours ago
- d6ae6a4 Fuse Gemma 4 global attention into tensor-core CUDA kernels by Google AI Edge · 33 hours ago
- 519999c Implement WebGPU binary, comparison, logical, and ternary (Select/SelectV2) operations for LiteRT Tensor API. by Ping Yu · 34 hours ago
- 7278fe8 Add LiteRT Quantizer (litert_quantizer) to the LiteRT repository. by Maria Lyubimtseva · 2 days ago
- d408614 Remove org_tensorflow and tf_workspace dependencies from LiteRT OSS by Terry Heo · 2 days ago chromium/8083
- b460105 Rewrite QuantizationTests.swift in XCTest and drop the swift-testing dep due to OSS rules_swift version mismatch. by Jun Jiang · 2 days ago
- 2588243 Internal change by Tenghui Zhu · 2 days ago
- b263ebb Require cl_arm_import_memory_android_hardware_buffer for AHWB <-> OpenCL interop. by Fengwu Yao · 2 days ago
- 769e75b Add LITERT_WITH_TENSORFLOW to test TensorFlow targets in CI by Terry Heo · 2 days ago
- 3445317 Add DMA-BUF and peak memory reporting to LiteRT tools and runner script for embedding models. by Andrew Zhang · 2 days ago
- f60efba Prevent redundant TFLite ArenaPlanner allocations for subgraph I/O by Terry Heo · 2 days ago
- c0575cf Add INT4 FullyConnected test coverage to LiteRT ATS. by Gerardo Carranza · 2 days ago
- 91b2c64 Implement WebGPU unary, activation, cast, and shape-aliasing operations for LiteRT Tensor API. by Ping Yu · 2 days ago
- b585783 Extract functions that can be reused to run with YNNPACK. by Quentin Khan · 2 days ago
- 5ee5bda Automated Code Change by Google AI Edge · 2 days ago
- c4d5081 Remove source_location shim now that absl::SourceLocation is open source. by Quentin Khan · 2 days ago
- 0b67fbc Add GQA head ratio (H_v vs H_k) support to CPU and GPU gated_delta_update kernels. by Google AI Edge · 2 days ago
- 321b969 Add a test in attention when the KV cache grows. by Quentin Khan · 2 days ago
- 3291071 Prepare tests to run over multiple backends. by Quentin Khan · 2 days ago
- 68a99d8 Extract buffer mapping logic that can be shared between XNNPACK and YNNPACK runners. by Quentin Khan · 2 days ago
- 165f86d Support broadcast-based grouped-query attention (GQA) during prefill in YNNPACK. by Volodymyr Kysenko · 2 days ago
- a0c165e Unify Flash-Decode SDPA across Metal and OpenCL with UCL wave-SIMD support. by Fengwu Yao · 2 days ago
- 75d54fd Add TensorFlow's third-party repositories to the LiteRT OSS shim by Terry Heo · 2 days ago
- 826c40d Protect XNNPACK runtime creation and destruction with workspace_mutex_. by Google AI Edge · 2 days ago
- e272cd5 Merge pull request #10032 from graham0824:dev/hungjuiw/fix-addn-test by Copybara-Service · 2 days ago
- 8a9cc74 Merge pull request #9665 from graham0824:dev/hungjuiw/qcom-target-backend by Copybara-Service · 2 days ago
- c00c771 [LiteRT][Qualcomm] Set HTP file read memory budget and release mmap pages after QNN context deserialization. by Weiyi Wang · 2 days ago
- b015746 Add `cint2_fp32_int4_e8m0_drq` fusion pass and reference `FullyConnected` kernel. by Majid Dadashi · 2 days ago
- 59097e4 Add `cint2_fp32_int4_e8m0_drq` fusion pass and reference `FullyConnected` kernel. by Majid Dadashi · 2 days ago
- e868789 Merge pull request #9480 from graham0824:dev/mingxiup/expand_dims_op by Copybara-Service · 2 days ago
- 3233ef5 Merge pull request #10188 from graham0824:dev/hungjuiw/transformation-in-compile by Copybara-Service · 2 days ago
- 950d251 Fix SdpaTransposed reference evaluation for bool masks and 32-aligned active KV length by Gerardo Carranza · 2 days ago
- 0159226 Migrate remaining ML Drift enum references in //third_party/odml to kPascalCase. by Juhyun Lee · 2 days ago
- acbb3a0 Add stand-in XLA and LLVM repositories for LiteRT OSS by Terry Heo · 2 days ago
- 3884c94 Add a stand-in TensorFlow repository for LiteRT OSS by Terry Heo · 3 days ago
- 4fe2408 Bump pinned XNNPACK commit to pick up qs8_qc4w neoni8mm GEMM kernels. by Changming Sun CA · 3 days ago
- dbc8e13 Package TensorFlow Lite CoreML and Metal delegates as standalone xcframeworks and expose CLiteRT_static in SwiftPM. by Jun Jiang · 3 days ago
- 76fe528 Make ml_drift delegate C++17 compliant. by Google AI Edge · 3 days ago
- c112443 Update readme of LiteRT github repo by Google AI Edge · 3 days ago
- b6a9c77 Optimize MoE GPU delegate graph for decode and top-k reduction. by Fengwu Yao · 3 days ago
- 143d165 This is an internal change. by Chunlei Niu · 3 days ago
- 34d6836 Migrate ML Drift enum references in //third_party/odml/litert/ml_drift to kPascalCase. by Juhyun Lee · 3 days ago
- 391c38f Fix heap-use-after-free in moe_experts_parser_test under ASAN. by Fengwu Yao · 3 days ago
- c7bc981 Add an internal only Kotlin API to control fallback to CPU in tests. by Chunlei Niu · 3 days ago
- 516e119 Migrate ML Drift enum references in //third_party/odml/litert/tensor to kPascalCase. by Juhyun Lee · 3 days ago
- 7f8b9a7 Automated Code Change by Google AI Edge · 3 days ago
- ee5819f Consolidate MoE builder utilities. by Raman Sarokin · 3 days ago
- 0750b18 Optimize SELECT_V2 inner loops by hoisting scalar broadcast values. by Dillon Sharlet · 3 days ago
- 3971f72 Add grouped-query SDPA kernel tests for the decomposed graph. by Fengwu Yao · 3 days ago
- 732dea9 Automated Code Change by Google AI Edge · 3 days ago
- 77a7a2c Harden SdpaTransposed ATS generator and fix MLDrift SDPA mask/softcap handling. by Gerardo Carranza · 3 days ago
- 6d59736 Fix SdpaTransposed ATS generator Q tensor layout and GPU GQA fallback. by Gerardo Carranza · 3 days ago
- efb191a Support blockwise int8 MoE expert scales in the XNNPACK delegate. by Fengwu Yao · 3 days ago
- 85d78cd Include QAIRT version in Qualcomm SDK package description. by Andrew Zhang · 3 days ago
- 58d4927 Use ucl::Init<float4> in the short_conv_step kernel. by Fengwu Yao · 3 days ago
- ec4de24 [LiteRT][MediaTek] Fix dangling pointer in GetDimensions, enum sanitizer traps, and SoC case sensitivity by Andrew Zhang · 3 days ago
- 873d8e5 Remove TensorFlow deps from strip_buffers and analyzer_wrapper by Terry Heo · 3 days ago
- adce58e Support blockwise-quantized MoE expert weights on the GPU delegate. by Fengwu Yao · 3 days ago